Delphi MethodResearch MethodsAcademic Writing

How to Analyse Delphi Study Results: Consensus Metrics and Reporting

5 min read

Finishing data collection in a Delphi study puts you in front of a specific analysis challenge: you have ratings from multiple rounds, a definition of consensus you stated in advance, and probably some items that reached agreement and others that didn't. How you handle each of these in analysis and reporting determines whether your results chapter holds up to scrutiny.

The Two Things You're Always Measuring in Delphi

Every Delphi study measures two separate constructs, and conflating them is one of the most common analysis mistakes. The first is consensus — whether the group has reached a sufficient level of agreement on an item. The second is stability — whether the distribution of ratings has stopped changing meaningfully between rounds. A study that reports only consensus without stability has technically not run a Delphi study; it's run a multi-round survey. Stability is what makes the iterative structure meaningful.

Choosing the Right Consensus Metric

The choice of consensus metric should be determined by your rating scale and your research question, and it should be made before data collection — not selected post-hoc to maximise the number of items reaching consensus.

Percentage Agreement

Percentage agreement calculates the proportion of panelists whose ratings fall within a defined range — typically within one or two points of the median on a Likert scale, or above a threshold value. It is intuitive, easy to report, and widely understood by non-specialist readers.

  • Typical threshold: 70–80% of panelists rating within the agreed range
  • Best for: Studies using relatively simple rating scales where the overall direction of consensus matters more than fine-grained distinctions
  • Weakness: Can overstate consensus when ratings cluster at opposite poles of the scale
  • Interquartile Range (IQR)

    The IQR measures dispersion by calculating the spread between the 25th and 75th percentile of the rating distribution. Narrow IQR indicates high agreement; wide IQR indicates disagreement. It is the standard metric in clinical Delphi research and is specifically required by several reporting frameworks including CREDES.

  • Typical threshold: IQR ≤1 or IQR ≤1.5 on a 9-point scale; IQR ≤1 on a 5-point scale
  • Best for: Studies using multi-point Likert scales where direction and magnitude of expert ratings are both informative
  • Weakness: Requires a minimum panel size to be meaningful — IQR calculated from 7 panelists is less stable than from 15
  • Level of Agreement Scores

    Some studies calculate more complex agreement indices such as Content Validity Index (CVI), common in nursing and healthcare instrument development. These are appropriate when the study has a specific psychometric purpose beyond general consensus — for example, validating a clinical measurement instrument.

    Calculating and Reporting Stability

    Stability is measured by comparing the rating distribution between consecutive rounds. The standard approach is to calculate the percentage of panelists who changed their rating by more than one point between rounds — a change rate below 15% is commonly used as a stability criterion, though this varies by field and should be stated in your protocol.

    Report stability separately from consensus in your results chapter. Items can be stable without having reached consensus — the group remains stably divided — and in those cases the stability result is itself a finding.

    Handling Non-Consensus Items

    Non-consensus items — items that did not reach your predefined threshold after all rounds — are findings, not failures. How you handle them in your results chapter is a methodological decision:

  • Report them as areas of expert disagreement — if the goal of your study was to map the expert opinion landscape rather than exclusively reach agreement, non-consensus items belong in your results with analysis of what the distribution reveals
  • Exclude them from final recommendations — if your study was generating clinical guidelines or policy recommendations, non-consensus items are typically excluded from the output while being reported transparently in the methods and results
  • Investigate the reasons for disagreement — if you collected open-ended reasoning between rounds, non-consensus items often have the richest qualitative material attached to them
  • Never omit non-consensus items from the results section without noting that they existed and what threshold they failed to meet.

    Round-by-Round Reporting

    A Delphi results chapter should report each round separately, not only the final outcome. Reviewers and examiners expect to see how the distribution changed across rounds. Minimum reporting per round:

  • Response rate (number of panelists who completed that round)
  • Distribution summary per item (median, IQR or percentage agreement, range)
  • Number of items reaching consensus
  • Stability measure where applicable
  • What the Final Results Section Should Look Like

    The most common and defensible structure for a Delphi results section:

  • Panel composition and response rates — who participated, how many completed each round, attrition by subgroup if relevant
  • Round 1 results — initial distributions, items removed or added, qualitative summary of emergent themes from open-ended responses
  • Round 2 (and 3 if applicable) results — distribution changes, stability assessment, items reaching and failing to reach consensus
  • Final consensus list — items that met your predefined threshold with their final distributions clearly reported
  • Non-consensus items — reported with final distributions and a brief note on the nature of the disagreement
  • A table showing each item with its Round 1 and final-round median, IQR, and consensus status is the clearest way to present this data and is standard in published Delphi papers across fields.

    Continue Learning in Delphi Academy

    Dive deeper into these topics with our comprehensive guides:

    Ready to Implement Your Delphi Study?

    Our AI-powered platform streamlines every step of the Delphi process—from expert recruitment to automated analysis and report generation.

    Related Articles

    Delphi MethodResearch Methods

    Real-Time Delphi vs. Traditional Delphi: How to Choose

    Two researchers can run methodologically sound Delphi studies and end up with completely different processes. The format choice matters more than most methods textbooks admit.

    4 min read
    How to Analyse Delphi Study Results: Consensus Metrics and Reporting | Durvey