This vignette will use measures of agreement, in particular, the Fleis-Cohen kappa, and the Goodman-Kruskal lambda to explore the proposed adequacy of a cognitively diagnostic assessment.
This assessment is based on an example found in [@Mislevy1995]. It is a test of language arts in
which there are four constructs being measured: Reading,
Writing, Speaking and Listening. Each
variable can tale on the possible values Advanced,
Intermediate or Novice.
There are four kinds of tasks:
There are 5 Reading, 5 Listening, 3 Writing and 3 Speaking questions on a form of the test. Therefore, the Q Matrix looks like:
| Reading | Writing | Speaking | Listening | |
|---|---|---|---|---|
| R1 | 1 | 0 | 0 | 0 |
| R2 | 1 | 0 | 0 | 0 |
| R3 | 1 | 0 | 0 | 0 |
| R4 | 1 | 0 | 0 | 0 |
| R5 | 1 | 0 | 0 | 0 |
| W6 | 1 | 1 | 0 | 0 |
| W7 | 1 | 1 | 0 | 0 |
| W8 | 1 | 1 | 0 | 0 |
| L9 | 0 | 0 | 0 | 1 |
| L10 | 0 | 0 | 0 | 1 |
| L11 | 0 | 0 | 0 | 1 |
| L12 | 0 | 0 | 0 | 1 |
| L13 | 0 | 0 | 0 | 1 |
| S14 | 1 | 0 | 1 | 1 |
| S15 | 1 | 0 | 1 | 1 |
| S16 | 1 | 0 | 1 | 1 |
Assuming that the parameters of the 5 item models are known, it is straightforward to simulate data from the assessment. This kind of simulation provides information about the adequacy of the data collection for classifying the students.
The simulation procedure is as follows:
Generate random proficiency profiles using the proficiency model.
Generate random item responses using the item (evidence) models.
Score the assessment. There are two variations:
Modal (MAP) Scores: The score assigns a category to each of the four constructs. (These are the columns named “mode.Reading” and similar.)
Expected (Probability) Scores: The score assigns a probability of high, medium or low to each individual. (These are the columns named “Reading.Novice”, “Reading.Intermediate”, and “Reading.Advanced”.)
The simulation study itself can be found in
vingette("SimulationStudies",package="RNetica").
The results are saved in this package in the data set
language16.
The simulation data has columns containing the “true” value for the
four proficiencies—“Reading”, “Writing”, “Speaking”, and “Listening”—and
for values for the estimates—“mode.Reading”, “mode.Writing”, &c. The
cross-tabulation of a true value and its corresponding estimate is known
as a confusion matrix. This is a matrix \(\mathbf{A}\), were \(a_{km}\) is the count of the number of
simulated cases for which first variable (simulated truth) is \(k\) and the second variable (MAP estimate)
is \(m\). This can be built in R using
the table function.[^2]
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 216 | 46 | 0 |
| Intermediate | 37 | 398 | 45 |
| Advanced | 0 | 41 | 217 |
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 134 | 98 | 5 |
| Intermediate | 43 | 375 | 91 |
| Advanced | 14 | 146 | 94 |
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 130 | 91 | 1 |
| Intermediate | 65 | 390 | 78 |
| Advanced | 2 | 76 | 167 |
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 232 | 14 | 0 |
| Intermediate | 29 | 437 | 46 |
| Advanced | 0 | 71 | 171 |
The model used in this simulation is a Bayesian network. Like many
kinds of classification models its output is the probability that the
subject is in each category. So the columns Reading.Novice,
Reading.Intermediate, and Reading.Advanced
represent the probability based on the observed evidence that the
simulee’s reading ability is in those three categories. This is called
the marginal distribution as it is one side of a big multi-way
table with all of the varaibles. The mode.Reading, which we
used as the score before, just tells which category has the highest
probability. The table below shows the “true” (simulated) reading
ability and estimated marginal probabilities for the first five
simulees.
kbl(language16[1:5,c("Reading","Reading.Novice",
"Reading.Intermediate","Reading.Advanced",
"mode.Reading")],
caption="Reading data from first five simulees.",
digits=3) |>
kable_classic()| Reading | Reading.Novice | Reading.Intermediate | Reading.Advanced | mode.Reading |
|---|---|---|---|---|
| Intermediate | 0.565 | 0.435 | 0.000 | Novice |
| Advanced | 0.000 | 0.069 | 0.931 | Advanced |
| Novice | 0.954 | 0.046 | 0.000 | Novice |
| Novice | 0.703 | 0.297 | 0.000 | Novice |
| Advanced | 0.000 | 0.521 | 0.479 | Intermediate |
Look at the first row of the data. The first simulee was randomly
assigned a reading ability of Intermediate and then the
data were generated based on this estimate. The classification of the
first simulee based on the 16 item form is a 56.5% chance of
Novice and a 43.5% chance of Intermediate. So,
the “modal” (highest probability) score is Novice.
When building the confusion matrix, each observation add a count of one to a cell based on the true and estimated value. So the first simulee contributes 1 towards the cell in the second row and first column in the table. In building the expected table, this value is split across the row according to the probabilities, so the first simulee adds .565 to the first cell in the second row, .435 to the second cell and nothing to the third cell.
The table below shows the calculations for the first five simulees.
The first row is the sum of the Reading.Novice,
Reading.Intermediate and Reading.Advanced
variable for simulees whose true Reading value is Novice,
the second for true Intermediate, and third row for true
Advanced.
The expected value of the confusion matrix, \(\overline{\mathbf{A}}\) is calculated as follows: let \(X_i\) be the value of the first (simulated truth) variable for the \(i\)th simulee, and let \(p_{im} = P(\hat{X_i}=m|e)\) be the estimated probability that \(X_i=m\) for the \(i\)th simulee is \(m\). Then \(\overline{a_{km}} = \sum_{i: X_i=k} p_{im}\).
In the running example, the first row (novice) is the
sum of all rows for which the true scores is novice. The
second row is the sum of the intermediate rows and the
third row advanced.
The function expTable() does this work. Note that it
expects the marginal probability to be in a number of columns marked
var.state, where var is the name
of the variable and state is the name of the state. If the data
uses a different naming convention, this can be expressed with the
argument pvecregex which is a regular expression with
special symbols <var> to be substituted with the
variable name and <state> to be substituted with the
state name.
The table below shows the expected matrix from the first five rows of the reading data.
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 1.658 | 0.565 | 0.000 |
| Intermediate | 0.342 | 0.435 | 0.589 |
| Advanced | 0.000 | 0.000 | 1.411 |
Based on this definition of the expected confusion matrix, here are the values for the four proficiency variables.
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 206.702 | 51.083 | 0.168 |
| Intermediate | 55.238 | 372.955 | 58.196 |
| Advanced | 0.060 | 55.962 | 199.636 |
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 125.296 | 85.976 | 27.997 |
| Intermediate | 86.891 | 284.228 | 134.581 |
| Advanced | 24.812 | 138.796 | 91.421 |
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 123.018 | 113.348 | 10.582 |
| Intermediate | 91.180 | 317.257 | 91.247 |
| Advanced | 7.802 | 102.395 | 143.171 |
|
Estimated
|
|||
|---|---|---|---|
| Novice | Intermediate | Advanced | |
| Simulated | |||
| Novice | 208.522 | 34.682 | 0.200 |
| Intermediate | 36.774 | 385.996 | 79.082 |
| Advanced | 0.704 | 91.321 | 162.718 |
Suppose two different classifiers (be they human or algorithic) are classifying the same set of individuals. The confusion matrix provides counts for how often they agree or disagree. In an unweighted agreement measure, agreement is binary: the classifiers had the same or different results. However, if the categories are ordered, off-by-one can be considered better than off-by-two. The weights specify how much better.
If we look at the diagonal of one of these confusion matrixes, this is the number of cases that exactly agree between the simulated truth and the estimate. The proportion is the agreement rate.
Obviously the closer to 1, the better. However, this will depend on the prevalence of the categories in the population. If almost all the cases are in one category, it is easy to get good classification accuracy. Cohen’s kappa and Goodman and Kruskals lambda are both adjustments that can be used.
The sum of the diagonal of the confusion matrix, \(\sum_k a_{kk}\), gives a count of how many
cases are exact agreements (in this case between the simulation and
estimation). Let \(N=\sum_k\sum_m
a_{km}\); then the agreement rate is \(\sum_k a_{kk}/N\). For the reading data
using the MAP scores, this is 831 out of 1000, so over 80% agreement.
The function accuracy() calculates the agreement rate.
| MAP | EAP | |
|---|---|---|
| Reading | 0.831 | 0.779 |
| Writing | 0.603 | 0.501 |
| Speaking | 0.687 | 0.583 |
| Listening | 0.840 | 0.757 |
Raw agreement can be easy to achieve if there is not much variability
in the population. For example, if 80% of the target population was
intermediate, a classifier that simply classified each respondent as
intermediate would achieve 80% accuracy, and one that
guessed intermediate randomly 80% of the time would achieve
at least 64% accuracy. For that reason, two adjusted agreement rates,
lambda and kappa, are often used.
There is a problem with the agreement rate: depending on the proportions of cases in the population, it could be very easy to get a good agreement rate. Suppose the question is whether or not a randomly chosen person has traveled beyond the Earth’s atmosphere. As there are only a few dozen astronauts and cosmonauts, just guessing that everybody has not been to space is likely to get 100% accuracy.
@Goodman1954 suggested a correction.
Pick the row which has the biggest proportion. For the Reading data,
this is Intermediate, there are 480 intermediate readers in
the sample, so simply classifying everybody as an intermediate would
correctly classify 48% of the sample.
Let \(p_{ij}\) be the proportion of cases where the “true” category was \(i\) and the estimated category was \(j\). Let \(p_{i+} = \sum_{j} p_{ij}\) be the row sums and \(\sum_{i} p_{ii}\) is the agreement. Lambda discounts the agreement by the biggest of the row sums.
\[ \lambda = \frac{\sum p_{ii} - \max_{i}
p_{i+}}{1 - \max_{i} p_{i+}}\] The function
gkLambda() will do this calculation. Here are the values
for the language test using both the MAP and expected agreements.
| MAP | EAP | |
|---|---|---|
| Reading | 0.675 | 0.570 |
| Writing | 0.191 | -0.010 |
| Speaking | 0.330 | 0.167 |
| Listening | 0.672 | 0.513 |
This is pretty good. There is a 67.5% improvement in classification
from the one-size-fits-all model of classifying everybody at
Intermediate.
Suppose the reading test was being used as a placement test.
If the test was not available, the only strategy would be a
one-size-fits-all one, assuming that all subjects are at the same level
of the variable, Intermediate. On the other hand, by giving
the reading test, the classification rate improves by 67%; this looks
like it might be useful.
Jacob Cohen (Flies, Levin & Paek, 2003) took a different approach which treats the two assignments of cases to categories more symmetrically. The idea is that these are two raters and the goal is to judge the extent of the agreement. Baseline here is to imagine two raters one of which assign categories randomly with probabilities \(a_{k+}/N\) and \(a_{+k}/N\). Then the expected number of random agreements is \(\sum_{k} a_{k+}a_{+k}/N\). So the agreement measure adjusted for random agreement is:
\[ \kappa = \frac{\sum_{k} a_{kk} - \sum_{k}a_{k+}a_{+k}/N}{N-\sum_{k}a_{k+}a_{+k}/N} \ .\]
Again this runs from -1 to 1. The function fcKappa
calculates Cohen’s kappa.
| MAP | EAP | |
|---|---|---|
| Reading | 0.733 | 0.651 |
| Writing | 0.329 | 0.197 |
| Speaking | 0.478 | 0.325 |
| Listening | 0.740 | 0.609 |
All three of kappa, lambda and raw agreement all assume that any missclassification is equally bad. Fleis suggest adding weights, were \(1 \ge w_{km} \ge 0\) is the desirability of classifying a subject who is \(k\) as \(m\). In this case, weighted agreement is \(\sum_k\sum_m w_{km} a_{km}/N\). The weighted versions of kappa and lambda are given by:
\[ \lambda = \frac{\sum\sum w_{km} a_{km} - \max_k \sum_m w_{km} a_{km}}{N - \max_k \sum_m w_{km} a_{km}} \ ;\]
\[ \kappa = \frac{\sum\sum w_{km} a_{km} - \sum_k \sum_m w_{km} a_{k+}a_{+m}/N}{N - \sum_k \sum_m w_{km} a_{k+}a_{+k}/N} \ .\]
There are three commonly uses cases:
Both linear and quadratic weight have increasing penalties for the number of categories of difference. So off-by-one has a lower penalty than off-by-two.
The accuracy, gkLambda and
fcKappa have both a weights argument where no
weights (“None”, default), “Linear” or “Quadratic” weights can be
selected, and a w argument where a custom weight matrix can
be entered.
wacc.tab <- data.frame(
None = sapply(cm,accuracy,weights="None"),
Linear = sapply(cm,accuracy,weights="Linear"),
Quadratic = sapply(cm,accuracy,weights="Quadratic"))
wlambda.tab <- data.frame(
None = sapply(cm,gkLambda,weights="None"),
Linear = sapply(cm,gkLambda,weights="Linear"),
Quadratic = sapply(cm,gkLambda,weights="Quadratic"))
wkappa.tab <- data.frame(
None = sapply(cm,fcKappa,weights="None"),
Linear = sapply(cm,fcKappa,weights="Linear"),
Quadratic = sapply(cm,fcKappa,weights="Quadratic")
)| None | Linear | Quadratic | |
|---|---|---|---|
| Reading | 0.831 | 0.915 | 0.958 |
| Writing | 0.603 | 0.792 | 0.886 |
| Speaking | 0.687 | 0.842 | 0.919 |
| Listening | 0.840 | 0.920 | 0.960 |
| None | Linear | Quadratic | |
|---|---|---|---|
| Reading | 0.675 | 0.675 | 0.675 |
| Writing | 0.191 | 0.153 | 0.075 |
| Speaking | 0.330 | 0.323 | 0.310 |
| Listening | 0.672 | 0.672 | 0.672 |
| None | Linear | Quadratic | |
|---|---|---|---|
| Reading | 0.733 | 0.780 | 0.837 |
| Writing | 0.329 | 0.393 | 0.479 |
| Speaking | 0.478 | 0.550 | 0.645 |
| Listening | 0.740 | 0.782 | 0.834 |