Delphi MethodResearch MethodsData AnalysisAcademic Writing

What Is Delphi Consensus? Metrics Explained

8 min read

Many Delphi studies report "consensus was reached" without ever defining what that means in numbers. Reviewers increasingly push back on this, because consensus in the Delphi method isn't a feeling — it's a statistical threshold that has to be defined before the study starts, measured the same way every round, and reported in full.

This guide breaks down the metrics actually used to measure Delphi consensus, where each one applies, the most common mistakes that get flagged in peer review, and how to decide on a threshold before you collect a single response.

What Counts as Consensus in a Delphi Study

There's no single universal definition of consensus across published Delphi research — that's well documented and somewhat uncomfortable for a method built on systematic rigor. The most common approach is percentage agreement, with 75% as the median threshold researchers select, though published thresholds range anywhere from 50% to 97%. Other studies define consensus through the interquartile range (IQR) of responses, and a smaller number use standard deviation or mean-based criteria.

What matters for your study isn't which definition is "correct" — there isn't one — but that you pick a definition, justify it, and apply it consistently and a priori, meaning before you've seen any data.

The Main Consensus Metrics

Percentage Agreement

This is the most intuitive metric: the share of panelists whose rating falls within a defined range, usually the same response category or within one point of the median. A common threshold is 70-75% agreement to declare an item has reached consensus. The tradeoff is that percentage agreement depends heavily on your scale length and the range you count as "agreeing" — a 5-point scale and a 9-point scale need different agreement bands to mean the same thing.

Interquartile Range (IQR)

IQR measures the spread of the middle 50% of responses — sort all ratings, find the value at the 25th and 75th percentile, and take the difference. A smaller IQR means panelists are clustered tightly around the same answer. On a 5-point Likert scale, an IQR of 1 or less is commonly treated as consensus; on a 9-point scale, a threshold of 2 or less is more typical. IQR is widely considered more statistically rigorous than percentage agreement because it doesn't depend on picking a specific category boundary.

Standard Deviation and Mean

Less common in Delphi reporting, but used when response data is treated as interval/ratio rather than ordinal — for example, numeric estimates rather than Likert ratings. A shrinking standard deviation across rounds signals convergence. This approach requires more caution: it assumes your scale behaves like a true numeric measure, which a 5-point Likert scale technically does not.

Stability: The Metric Reviewers Often Miss

Consensus and stability are not the same thing, and conflating them is one of the more common reporting gaps. Consensus describes how tightly panelists agree at a single point in time. Stability describes whether that agreement holds steady between rounds, rather than shifting back and forth as panelists keep changing their ratings.

A panel can hit your consensus threshold in Round 2 purely by chance, then drift apart again in Round 3. Without a stability check — typically defined as the percentage of items whose rating doesn't change by more than one scale point between consecutive rounds — you can't tell the difference between genuine convergence and statistical noise. Define your stability criterion alongside your consensus threshold, before the study starts, not after you've seen how the data is trending.

Choosing and Reporting Your Threshold Before You Start

The single most important decision in this entire process is timing: your consensus and stability thresholds need to be set before Round 1 launches, not selected after looking at how the data came out. Reviewers are specifically trained to look for this now, and a threshold that conveniently matches your actual results looks like post-hoc rationalization even when it isn't.

In your methods section, state explicitly:

  • The exact metric you're using (percentage agreement, IQR, or both)
  • The numeric threshold that constitutes consensus for that metric
  • The scale length the threshold is calibrated to (a 5-point and 9-point scale need different numbers)
  • Your stability criterion and how many consecutive rounds it must hold for
  • What happens to items that never reach consensus — dropped, flagged as "no consensus," or carried forward with a hard round limit
  • Common Mistakes That Undermine Consensus Claims

  • Defining the threshold after seeing the data. This is the fastest way to get a methodology section flagged in review.
  • Reporting only the final round's statistics. Reviewers want to see how consensus evolved round by round, not just the end state.
  • Treating "majority voted this way" as statistical consensus. A simple majority is not the same as meeting a pre-defined IQR or agreement threshold.
  • Skipping the stability check entirely. A single round of tight agreement can be coincidental — stability is what confirms it's real.
  • Using an agreement percentage without specifying the scale it's calibrated to. "70% agreement" means something different on a 4-point scale than a 9-point scale.
  • How Durvey Calculates Consensus Automatically

    Durvey calculates median, IQR, and percentage agreement per item in real time as responses come in, rather than requiring you to export data and recompute statistics by hand after every round. Convergence and stability across rounds are tracked the same way, so you can see exactly when — and whether — a genuine, stable consensus has been reached. The export format maps directly to what a CREDES-aligned methods and results section needs, including round-by-round breakdowns rather than just a final snapshot.

    A Worked Example

    A panel of 20 experts rates an item on a 9-point importance scale. In Round 1, the median is 6 with an IQR of 3 — wide disagreement. After structured feedback showing the group distribution, Round 2 produces a median of 7 with an IQR of 1.5, which is closer to the pre-defined consensus threshold of IQR ≤ 2 but not stable yet, since 30% of panelists shifted their rating by more than one point from Round 1. Round 3 produces a median of 7 with an IQR of 1, and fewer than 10% of panelists move by more than one point from Round 2 — meeting both the consensus threshold and the stability criterion at the same time. That's the point at which you can report consensus was reached, with the numbers to back it up.

    Continue Learning in Delphi Academy

    Dive deeper into these topics with our comprehensive guides:

    Ready to Implement Your Delphi Study?

    Our AI-powered platform streamlines every step of the Delphi process—from expert recruitment to automated analysis and report generation.

    Related Articles

    What Is Delphi Consensus? Metrics Explained | Durvey