Item Analysis Made Simple: Formula, Interpretation Table, and Worked Examples

Teacher Guide June 14, 2026 · 5 min read

You’ve checked the papers. Now comes the part that turns a pile of scores into better teaching: item analysis — finding out which questions your class struggled with, and why. This guide covers the two core formulas with worked examples, the interpretation tables, distractor analysis, the sex-disaggregated reporting DepEd increasingly asks for, and — most importantly — what to actually do with the results.

The big idea: a low class average doesn’t tell you what to fix. Item analysis does. It separates “the learners didn’t study” from “item 7 was miskeyed” from “I never actually taught that competency” — three very different problems with three very different solutions.

1. The difficulty index (the one you’ll use most)

For each test item, the difficulty index answers: how many learners got it right?

Difficulty Index (p) = Number of learners who answered correctly ÷ Total number of learners

Example: 27 of 45 learners got item 5 right → p = 27 ÷ 45 = 0.60.

Despite the name, a higher index means an easier item. (Think of it as the “pass rate” for that single question.)

Difficulty interpretation table

Difficulty index Interpretation Action
0.81 – 1.00 Very easy Fine as a confidence-builder; too many = test too easy
0.61 – 0.80 Easy Keep
0.41 – 0.60 Moderately difficult Ideal range — keep
0.21 – 0.40 Difficult Review the item and reteach the skill
0.00 – 0.20 Very difficult Red flag — check for a miskey, confusing stem, or untaught content

A healthy test has most items between 0.41 and 0.80. A test where every item sits above 0.85 isn’t “a smart class” — it’s a test that didn’t discriminate, and it tells you nothing about who needs help.

2. The discrimination index (does the item separate strong from weak?)

A good item is one that strong learners tend to get right and struggling learners tend to get wrong. The discrimination index measures that. Rank papers by total score, take the top 27% and bottom 27% of learners, then:

Discrimination Index (D) = (Correct in upper group − Correct in lower group) ÷ Number of learners per group

D Interpretation
0.40 and up Excellent item
0.30 – 0.39 Good
0.20 – 0.29 Marginal — revise
Below 0.20 Poor — discard or rewrite
Negative Strong learners got it wrong more than weak ones — almost always a miskeyed answer. Check your key first!

A negative discrimination index is the most actionable signal in all of item analysis: it nearly always means the answer key is wrong, not the learners.

3. Putting both together: a worked example

40 learners take a test. Looking at item 12:

  • 14 learners answered correctly → p = 14 ÷ 40 = 0.35Difficult
  • Upper 27% (about 11 learners): 8 correct
  • Lower 27% (about 11 learners): 3 correct
  • D = (8 − 3) ÷ 11 = 0.45Excellent discrimination

Verdict: a hard but fair item. The content was genuinely difficult (low p), but the item works — strong learners got it, weak learners didn’t (high D). Action: reteach the skill; keep the item. Contrast that with an item at p = 0.35 and D = 0.05: same difficulty, but now nobody is reliably getting it — that item is broken, not hard.

4. Distractor analysis (for multiple choice)

Beyond the indices, look at which wrong options learners chose. For each item, tally how many picked A, B, C, D:

Option Count Note
A 3 weak distractor — barely chosen
B ✓ 14 the key
C 18 a stronger pull than the answer — investigate!
D 5 working distractor

When a distractor attracts more learners than the correct answer (option C above), one of two things is true: it’s a genuine misconception worth reteaching, or the item is ambiguous and C is defensible. Both are worth knowing. A distractor that nobody picks (option A) is dead weight — replace it next time.

5. Sex-disaggregated item analysis

DepEd reporting increasingly asks for sex-disaggregated data — results broken down by male and female learners. For item analysis, that means computing the difficulty index per item two more times: once for boys, once for girls.

Item All (p) Male (p) Female (p) Note
3 0.62 0.61 0.63 No gap
7 0.48 0.31 0.66 Large gap — investigate item context

A consistent gap on certain item types is genuinely useful for your LAC session — that’s the point of the requirement, not just compliance. (We cover the full reporting side in our sex-disaggregated data guide.)

6. What to actually do with the results

This is where item analysis earns its keep:

  • p below 0.40 across a cluster of items from one competency → that competency needs reteaching, not just item rewrites. The test surfaced a teaching gap.
  • One item at p < 0.20 while the rest are fine → suspect the item: ambiguous stem, two defensible answers, or a miskey.
  • Negative discriminationfix the answer key before anything else.
  • A distractor outpulling the key → reteach the misconception, or rewrite the option.
  • Keep the good items. Items in the ideal ranges (p 0.41–0.80, D ≥ 0.30) belong in your item bank for next year’s parallel test — you’re slowly building a tested, reliable question pool.

Frequently asked questions

How many learners do I need for reliable item analysis?
The indices stabilize with larger groups, but even one class of 30–45 surfaces the actionable patterns — miskeys, untaught content, gaps. Don’t wait for a big sample to start.

Do I analyze essay items?
The difficulty and discrimination indices are designed for objective items. For essays, analyze the average rubric score per criterion instead — it tells you which criterion the class struggled with.

Is the 27% rule mandatory?
It’s the classical convention (it maximizes the contrast between groups). Some schools use top/bottom 25% or 33% — any consistent split works for classroom purposes.

Should I drop items with poor discrimination from this test’s grades?
That’s a judgment call and often a school-policy question. A clearly miskeyed item is usually corrected for everyone; a merely “hard” item that’s otherwise sound is typically kept. Document your reasoning.

How often should I do this?
At minimum for every summative test and term exam — ideally before you finalize grades, so a broken item never costs a learner unfairly.


Skip the spreadsheet entirely

Skoolari’s Auto-Checker does all of this automatically: photograph your learners’ answer sheets with your phone, review the AI’s checking item by item, and it produces the score sheet plus a complete item analysis — per-item difficulty, sex-disaggregated correct counts for all, male, and female learners — downloadable as Excel. What used to take an evening takes a coffee break.

👉 Try Skoolari free — no credit card needed

The skoolari Team
Skoolari is an AI-powered class record & assessment system for Filipino teachers — MATATAG-ready records, AI test generation, and photo auto-checking. Try it free →

Ready to save hours every week?

Join teachers who are generating better assessments, grading faster, and keeping cleaner records — all in one place.

No setup required. Works in any modern browser.