Attribute Agreement Analysis evaluates the consistency of a measurement system when results are categorical rather than continuous.

Examples include:

  • pass/fail;
  • good/bad;
  • defect type;
  • acceptable/unacceptable;
  • severity category.

The method asks whether appraisers make the same decision repeatedly and whether different appraisers agree with one another.

Use it when judgment is part of measurement

Attribute inspection may depend on:

  • visual appearance;
  • sound;
  • fit;
  • document classification;
  • defect coding.

If people interpret the criteria differently, the resulting quality data can be unreliable.

Measurement System Analysis provides the broader framework for understanding variation created by the measurement process.

Evaluate repeatability

Repeatability asks:

Does the same appraiser classify the same item the same way when it is presented again?

If an inspector calls the same part acceptable in one trial and unacceptable in another, the method lacks repeatability.

Repeated trials should be arranged so the appraiser does not simply remember the earlier decision.

Evaluate reproducibility

Reproducibility asks:

Do different appraisers classify the same item the same way?

Disagreement may indicate:

  • vague criteria;
  • weak examples;
  • inconsistent lighting;
  • poor training;
  • difficult boundary conditions.

The purpose is to improve the measurement system, not rank inspectors.

Compare with a reference where possible

A reference or expert standard allows the study to ask:

Are appraisers agreeing with the correct classification?

High agreement between appraisers is not enough if they are consistently making the wrong decision.

The reference should be credible and defined before the study.

Include boundary samples

A study containing only obvious good and obvious bad samples may overstate agreement.

Include realistic difficult cases near the decision boundary.

These samples reveal whether the operational criteria are sufficiently clear.

Operational Definition helps convert vague terms such as “acceptable surface” into repeatable classification rules.

Review false accepts and false rejects

Two important errors are:

  • false accept: bad item classified as good;
  • false reject: good item classified as bad.

The consequence of these errors may be different.

A safety-critical false accept can matter more than a false reject that creates extra inspection cost.

Use the findings to improve the method

Possible improvements include:

  • clearer defect definitions;
  • visual standards;
  • better lighting;
  • improved fixtures;
  • training;
  • automated measurement.

Quality at the Source is stronger when the people making quality decisions have a measurement method they can apply consistently.

Revalidate after significant change

Repeat the analysis when there is a meaningful change in:

  • product;
  • defect criteria;
  • inspection equipment;
  • workforce;
  • customer standard.

The measurement system should remain suitable for the current process.

Common mistakes

Testing only obvious samples, using too few repeat trials, allowing appraisers to remember prior classifications, assuming agreement equals correctness, blaming inspectors instead of improving criteria, and failing to reassess after standards change are common mistakes.

Practical sequence

  1. define the attribute decision.
  2. establish a credible reference.
  3. select representative samples.
  4. include boundary cases.
  5. randomize sample presentation.
  6. have each appraiser classify samples repeatedly.
  7. evaluate within-appraiser agreement.
  8. evaluate between-appraiser agreement.
  9. compare with the reference.
  10. improve the method and repeat the study when needed.

The practical lesson

Attribute Agreement Analysis verifies whether categorical inspection decisions can be trusted.

Reliable quality data requires people to make the same decision consistently and, where possible, the correct decision.