Unofficial guide · not affiliated with the NCBE, the ABA, or any state bar ← SI Bar Exam · Exam Admin · Academy

Eli Item-Writing Academy

Academy › Application & Calibration

Difficulty Calibration: Estimating Tiers Honestly

About 35 minutes · Academy module: Difficulty Calibration: Estimating Tiers Honestly.

Learning goals

  • Explain what "estimated difficulty" means — and what it doesn't.
  • Apply the feature heuristic (stem length, options, fact-pattern density, multi-step phrasing) consistently.
  • Know when an estimate must give way to response data.

Lesson 1 — Estimates, Not Measurements

Every difficulty label in this system is an estimate from item features — never a calibrated psychometric value. The honest copy rule: UI and reports say "estimated difficulty," never "difficulty." Estimates exist to seed adaptive practice and balance exam forms; they are replaced by real response data (p-values) as soon as students answer the item.

Lesson 2 — The Feature Heuristic

The sidecar (see `js/difficulty.js`) scores observable features: longer stems cost more working memory; five options cost more than four; a fact pattern dense with legal detail costs more than a clean hypothetical; multi-step phrasing ("best explains," "most likely," "best defense") signals layered reasoning. Score 0–1 → easy; 2–3 → medium; 4+ → hard. The heuristic is documented, deterministic, and auditable — anyone can recompute any tier by hand.

Lesson 3 — Calibrate Yourself Against Data

Once an item has ~30+ responses, compare its estimated tier to its p-value. Systematic mismatches teach you about your own writing: if your "easy" items keep landing at p = 0.45, you're underestimating your distractors' pull. Keep a calibration log — it's the fastest way to become a better estimator.

Lesson 4 — When Estimates Must Not Be Used

Never use estimated tiers for high-stakes decisions: pass/fail cut scores, grades, or "readiness" verdicts. Estimates seed practice and balance forms; only response data plus faculty judgment make consequential calls. The system's flags are review signals, never verdicts.

Spot-the-flaw drill

Flaw: presenting an estimate as a measurement, and worse, as a readiness signal. The l

presenting an estimate as a measurement, and worse, as a readiness signal. The label was assigned by heuristic before a single student answered the item.

Correction

The honest framing: "estimated hard — long fact pattern, multi-step call." Readiness is a property of students, estimated from response data across many items, never of one item's label.

Model rewrite

— how the item should be presented:

Estimated difficulty: hard (long fact pattern, multi-step reasoning). Estimates come from item features, not student data — they'll update as responses come in.

Drill key

The label is the same; the honesty around it is the fix. Estimates guide practice; they never certify readiness.

Sources

  • Downing, S. M., & Haladyna, T. M. (Eds.). (2006). Handbook of test development. Lawrence Erlbaum Associates. (On the distinction between judgmental and empirical difficulty.)
  • Ebel, R. L. (1972). Essentials of educational measurement. (On why pre-administration difficulty judgments are estimates.)

Finished this module?

Recording it adds the module to your Academy completion record on this device — module, track, level, and date, ready to download from your account page for faculty-development documentation.