How do we estimate the EVA?
Students’ exam results depend on many factors independent of the school, most of all on prior students’ achievements. If it is so, exam results may be in many cases a highly misleading indicator of whether a given school teaches well. What is a much better teaching effectiveness indicator are measures of progress made by students while studying in a given school. In a statistically proper manner, the progress can be identified by the educational value added method. The EVA method is a set of statistical techniques that enable measurement of the school’s input into learning results. To be able to apply it, we need the results of two school achievements measurements: at the beginning of studying in a given school and after its completion.
In the EVA indicator estimation model for schools, we take into account information on exam results and, in addition, on the gender of the student and dyslexia. The conceptual model of estimation of EVA indicators for schools leading to the upper secondary school (matura) exam on the example of a general upper secondary school can be presented as follows.

Chart 1. Conceptual model of estimating the EVA for an LO
A model for technical upper secondary school would look analogously.
The Polish examination system provides data for calculation of the EVA for lower secondary schools (the sixth grade primary school test as the input measure, the lower secondary school examination as the output measure) and upper secondary schools (lower secondary school examination as the input measure, and the upper secondary school examination as the output measure).
The EVA indicator for a general upper secondary school reveals how high/low matura exam results were achieved by its leavers in comparison to the students with analogous lower secondary school exam results. Similarly, the EVA indicator for technical upper secondary schools shows how high/low matura exam results were achieved by its leavers in comparison to the students of technical upper secondary schools in Poland with statistical control over the results of the lower secondary school exam. The EVA indicators are of relative nature and do not provide information on absolute progress in learning. They serve to compare schools. At the national scale, the EVA indicator is, ex definitione, equal to zero. An additional value of the EVA indicates an above-average effectiveness of teaching at the national scale, while a negative value a lower than average effectiveness.
The most frequently applied synthetic measure that describes the teaching outcomes is the arithmetic mean of exam results. The mean value as an indicator of results has its advantages, but also some serious disadvantages. One advantage is the possibility to describe the teaching outcomes in a given school by means of one figure. It is a disadvantage, on the other hand, that a significant proportion of information is lost and there is high uncertainty of reasoning resulting from a different shape of distribution of exam results in specific schools. For instance, if all students in a given school obtain results close to the arithmetic mean, the mean value is a good measure of the teaching outcomes. If, however, most of the students in a given school obtains results much higher or lower than the average (high variability of results) or the distribution is not symmetrical, the average is a very imperfect indicator of teaching outcomes. Despite the weakness of the mean exam result as a measure of teaching outcomes, it is used in our presentation because of its synthetic nature. Since both a measure of teaching outcomes and teaching effectiveness (EVA) are presented in one graph, it is this quality of the mean that determines it use. In order to deepen the analyses of teaching outcomes for the purposes of intraschool evaluation, other statistical tools need to be used, such as distribution of raw results on a stanine scale.
It must be stressed that exam indicators, both of the final result and the EVA, are estimated separately for the populations of general and technical upper secondary schools. It means that the indicators for general and technical upper secondary schools are not comparable.
Why three-year indicators?
In estimating the indicators we use the results from three subsequent years. Therefore, we talk about three-year indicators. The way to count three-year indicators is well demonstrated in the following graph.

Each year, data from the oldest previously included class are removed and current results are added.
The calculation of three-year indicators is justified for three reasons. The most important is the fact that exam results are impacted by measurement uncertainty. The greater the pool of data at our disposal, the lesser the impact of uncertainty on estimation of indicators for a school. Using results from three consecutive years, we obtain three times as much data, and a large body of information means greater precision of estimates. School evaluation is, therefore, not dependent on temporary changes, a single “statistical oddity” increasing or decreasing the value of indicators.
Yet three-year indicators also have a disadvantage – they are less sensitive to actual short-term changes. Therefore, a school should not rest with analysis of long-term indicators, and intraschool analyses should include calculation, with the use of the EVA Calculator (see: http://ewd.edu.pl) annual measures.
What data are used to estimate indicators?
To calculate the presented evaluatory exam indicators, exam data for three cohorts of school leavers are used: 2010 through 2012. Long-term exam indicators are less exposed to random variations and, therefore, are a more reliable measure of teaching outcomes and effectiveness in specific schools.
For each school, evaluatory exam indicators are calculated for four areas of teaching. The first one, humanist, covers Polish, History and Social Sciences. The second area covers the results of the exam in Polish. The third subject area is Maths and Sciences, covering Maths, IT, Physics, Biology and Geography. The fourth area is formed by Maths, separated for a distinct indicator. In estimating exam indicators, information on performance of exam tasks, both on the basic and advanced levels are used.
The achievements of students at the threshold of the general or technical upper secondary school are determined on the basis of the data from the lower secondary school exam in the standard version for those taking the exam on the major session, in April. When calculating EVA matura indicators in the scope of the humanities, data from the humanities part of the lower secondary school exam are used, while for calculation of EVA indicators for maths and sciences, information from the analogous part of the lower secondary school exam is used.
The following chart summarises the EVA indicator estimation model.

Chart 2. Structure of evaluatory exam indicators
Both the average matura exam result and the EVA in a given area are presented on a scale normalised for the country – separately for general and technical upper secondary schools – on a standard scale with a mean of 100 and deviation of 15. A description of the scale can be found further below in the informational part.
On what scale are the results presented?
For estimation of evaluatory exam indicators, information on the level of performance for all tasks solved by the student taking the matura exam within a given study area for both level of the exam. When, for instance, the student wrote only the exam in Maths in the area of Maths and Sciences at the basic level, the level of his or her skills in this area is estimated only on the basis of a rather limited information base, that is the data on how he or she dealt with the Maths tasks at the basic level. If he or she also wrote the advanced level, we take into account additionally information on how he or she dealt with the tasks at this level. When the student also took exams in other subjects of this area, we take more information into consideration. Thanks to modern test result scaling methods (the so-called IRT method and selection method), we can estimate on one scale the level of skills of students, despite the fact that they did not solve the same set of tasks. With application of this method, the result obtained by the student taking the matura exam does not depend on the number of tasks performed, but on how he or she coped with tasks of various difficulty, which he or she dealt with during the exam effort.
Owing to the properties of the applied IRT scaling methods, the obtained estimates of the students’ levels of skills are the best calculation of indicators of teaching outcomes which can be derived from exam data. The application of those indicators for individual assessment of students’ achievements would be problematic, but those indicators make the best use of exam data for evaluation of teaching outcomes.
The indicators of the student skill level obtained for each of the four distinguished subject areas are converted into a standardised scale with a national mean of 100 and standard deviation of 15. The standardisation procedure is carried out separately for students of general and technical upper secondary schools, in each of the four aforementioned subject area indicator.
The graph below presents the assumed percentage distribution of results.

Thanks to the properties of normal distribution, the result presented on a standardised scale has a set quantitative interpretation, that is one can say what a given result looks like against the reference population. The graph presents percentage values for selected points of the scale. Thus, a result of 130 means that only slightly more than 2% of exam takers in a given population obtained a better result, while a result of 85 indicates that almost 16% obtained a lower result.
Scaled upper secondary school exam results are used both for presentation of the mean matura exam result for a given general or technical upper secondary school, as well as presentation of the value of the educational value added indicator. In the case of the EVA indicator, the scale has its centre in point 0. For example, a school’s EVA score equal to 5 points means that, in a given subject area, the matura exam result of an average student of that school was 5 points higher (on a standard scale, that is higher by one-third of the standard deviation) than it would result from his or her result at the lower secondary school exam.


