Unofficial guide · not affiliated with the NCBE, the ABA, or any state bar ← SI Bar Exam · Exam Admin · Academy

Eli Item-Writing Academy

Academy › Quality at Scale

Item Analysis: Reading What the Numbers Say

About 40 minutes · Academy module: Item Analysis: Reading What the Numbers Say.

Learning goals

  • Read p-values, discrimination indices, and distractor pull tables.
  • Distinguish "hard item" from "broken item" in the data.
  • Turn analysis into revision decisions — and know when to retire an item.

Lesson 1 — The P-Value (Difficulty)

The p-value is the proportion of examinees answering correctly. 0.30–0.80 is the informative band for most purposes. Below 0.30 the item may be too hard or flawed; above 0.90 it contributes little information (though easy items have their place early in a practice set for confidence).

Lesson 2 — Discrimination: Does The Item Know Its Audience?

The discrimination index (upper–lower 27%) and the point-biserial correlation ask the same question: do high scorers get this right more often than low scorers? Positive and healthy (point-biserial above ~0.30) means the item measures what the exam measures. Near zero means noise. Negative means the item is actively mismeasuring — strong students are drawn to a distractor — and it goes to faculty review immediately.

Lesson 3 — Distractor Pull Tables

For each option: what proportion chose it, and what was their average total score? Healthy pattern: the key pulls the top, each distractor pulls some of the middle, none pulls the top harder than the key. Pathologies: a dead distractor (near-zero pull — replace it), a seductive distractor (pulls top scorers — possible miskey or second defensible answer), and the flat item (all options pull evenly — the stem is probably unfocused).

Lesson 4 — From Numbers To Decisions

Data suggests; faculty decides. The revision menu: reword (a distractor is ambiguous), replace (a distractor is dead), rekey (the data plus re-solving shows the key was wrong — rare, documented, and never silent), retire (the item is unfixable or off-spec). Every change is logged with the reason. Items are never auto-deleted or auto-revised by the system — the analytics flags are review signals only.

Spot-the-flaw drill

Flaw:

Correction

The discrimination is negative: the students choosing B average 81% overall while those choosing the key average 62%. Strong students are systematically drawn to B. This item is mismeasuring — it goes to faculty review, not to the bank.

Model rewrite

— the review decision:

Flag for faculty review: negative discrimination (point-biserial −0.16). Distractor B pulls high scorers (avg 81%) harder than the key (avg 62%). Action: re-solve cold with two independent solvers; if B is defensible, rekey or retire — do not ship as-is.

Drill key

Hard ≠ good. A hard item that discriminates positively is valuable; a hard item with negative discrimination is broken. The numbers tell you which — if you read all of them.

Sources

  • Crocker, L., & Algina, J. (1986). Introduction to classical and modern test theory. Holt, Rinehart & Winston. (The standard reference for p-values, discrimination, and distractor analysis.)
  • Ebel, R. L., & Frisbie, D. A. (1991). Essentials of educational measurement (5th ed.). Prentice Hall.

Finished this module?

Recording it adds the module to your Academy completion record on this device — module, track, level, and date, ready to download from your account page for faculty-development documentation.