How to Analyse Delphi Study Results: Consensus Metrics and Reporting
Finishing data collection in a Delphi study puts you in front of a specific analysis challenge: you have ratings from multiple rounds, a definition of consensus you stated in advance, and probably some items that reached agreement and others that didn't. How you handle each of these in analysis and reporting determines whether your results chapter holds up to scrutiny.
The Two Things You're Always Measuring in Delphi
Every Delphi study measures two separate constructs, and conflating them is one of the most common analysis mistakes. The first is consensus — whether the group has reached a sufficient level of agreement on an item. The second is stability — whether the distribution of ratings has stopped changing meaningfully between rounds. A study that reports only consensus without stability has technically not run a Delphi study; it's run a multi-round survey. Stability is what makes the iterative structure meaningful.
Choosing the Right Consensus Metric
The choice of consensus metric should be determined by your rating scale and your research question, and it should be made before data collection — not selected post-hoc to maximise the number of items reaching consensus.
Percentage Agreement
Percentage agreement calculates the proportion of panelists whose ratings fall within a defined range — typically within one or two points of the median on a Likert scale, or above a threshold value. It is intuitive, easy to report, and widely understood by non-specialist readers.
Interquartile Range (IQR)
The IQR measures dispersion by calculating the spread between the 25th and 75th percentile of the rating distribution. Narrow IQR indicates high agreement; wide IQR indicates disagreement. It is the standard metric in clinical Delphi research and is specifically required by several reporting frameworks including CREDES.
Level of Agreement Scores
Some studies calculate more complex agreement indices such as Content Validity Index (CVI), common in nursing and healthcare instrument development. These are appropriate when the study has a specific psychometric purpose beyond general consensus — for example, validating a clinical measurement instrument.
Calculating and Reporting Stability
Stability is measured by comparing the rating distribution between consecutive rounds. The standard approach is to calculate the percentage of panelists who changed their rating by more than one point between rounds — a change rate below 15% is commonly used as a stability criterion, though this varies by field and should be stated in your protocol.
Report stability separately from consensus in your results chapter. Items can be stable without having reached consensus — the group remains stably divided — and in those cases the stability result is itself a finding.
Handling Non-Consensus Items
Non-consensus items — items that did not reach your predefined threshold after all rounds — are findings, not failures. How you handle them in your results chapter is a methodological decision:
Never omit non-consensus items from the results section without noting that they existed and what threshold they failed to meet.
Round-by-Round Reporting
A Delphi results chapter should report each round separately, not only the final outcome. Reviewers and examiners expect to see how the distribution changed across rounds. Minimum reporting per round:
What the Final Results Section Should Look Like
The most common and defensible structure for a Delphi results section:
A table showing each item with its Round 1 and final-round median, IQR, and consensus status is the clearest way to present this data and is standard in published Delphi papers across fields.
Continue Learning in Delphi Academy
Dive deeper into these topics with our comprehensive guides:
Ready to Implement Your Delphi Study?
Our AI-powered platform streamlines every step of the Delphi process—from expert recruitment to automated analysis and report generation.
Related Articles
Delphi Panel Size: How Many Experts Do You Actually Need?
The most common question before starting a Delphi study has no universally agreed answer — but the evidence points clearly in one direction.
Real-Time Delphi vs. Traditional Delphi: How to Choose
Two researchers can run methodologically sound Delphi studies and end up with completely different processes. The format choice matters more than most methods textbooks admit.