Skip to main content

Intercoder agreement in qualitative analysis

When more than one person codes the same data, it is natural to ask how far their decisions overlap; yet the question does not mean the same thing in every qualitative approach. This guide brings together the debate, common measures, a small worked calculation and practical steps.

Prepared by:
YouReply Qualitative content team
Published:
Last updated:
{minutes} min read
10 min read

Reliability and agreement are not the same

Intercoder agreement describes the extent to which two or more coders assign the same codes to the same units of data. Intercoder reliability is a more ambitious concept: it aims to show that the coding process is consistent enough to produce similar results when different people work independently. Agreement is an observation; reliability is an inference drawn from that observation about the process itself.

In their work on coding semistructured interviews, Campbell et al. (2013) draw a similar distinction: the overlap between decisions coders make independently is different from the agreement reached after disagreements have been discussed and reconciled. Reconciliation can improve the quality of the analysis, but it should not be reported as evidence of independent reliability.

Key concepts
ConceptWhat it asksTypical evidence
Intercoder agreementDid coders assign the same code to the same unit?Percent agreement, agreement table
Intercoder reliabilityIs agreement beyond what chance would produce, and is the process reproducible?Chance-corrected coefficients (kappa, alpha)
Negotiated agreementWas a shared decision reached after discussion?Records of disagreement meetings, codebook revisions

Does it fit qualitative paradigms?

O'Connor and Joffe (2020) show that the place of intercoder reliability in qualitative research is contested and that both sides have serious arguments. Supporters point out that measuring it increases systematicity and transparency, forces a structured team discussion about what codes mean and makes findings easier to communicate to readers from other disciplines. Critics argue that assuming a single “correct” coding sits uneasily with paradigms that treat meaning as dependent on context and interpreter. The authors recommend that the decision to measure agreement be made, and justified, in line with the study's aims and epistemological position.

Braun and Clarke (2019) take a clear position in this debate. In reflexive thematic analysis, coding is a flexible, evolving process that uses the researcher's subjectivity as a resource; there is no single correct coding to arrive at. Coding reliability measures therefore do not fit the logic of reflexive thematic analysis. The authors treat thematic analysis not as one method but as a family that includes coding reliability, codebook and reflexive approaches. Which type you work in largely determines whether measuring agreement is meaningful.

When does intercoder agreement make sense?

Measuring agreement is most useful when codes are defined in advance or early on and more than one person has to apply the same codebook. MacQueen et al. (1998) treat structured codebook development, together with regular checks on whether coders apply codes consistently, as part of improving the codebook in team-based qualitative analysis.

  • Codebook-based, team-based studies: You need to see whether different people understand and apply the same definitions in the same way.
  • Large datasets: When the coding workload is divided, decisions made by different coders must be combinable.
  • Applied research: Evaluation, policy or internal reports may be expected to show that findings were produced systematically.
  • Subsequent quantitative steps: If code frequencies will be compared across groups or passed on to statistical analysis, coding consistency directly affects the results.

In studies led by a single researcher, aimed at interpretive depth or taking a reflexive approach, agreement measures are generally not expected. Whatever you decide, stating the rationale clearly in your methods section helps readers judge your study by the right criteria.

The unitization problem: what are we comparing?

To calculate agreement, both coders must have made decisions about the same units. That is easy with naturally segmented data such as open-ended survey responses. In semistructured interviews, however, text does not divide itself into units: one coder may mark three sentences while the other marks the whole paragraph. Campbell et al. (2013) discuss this as the unitization problem; when unit boundaries differ, calculated agreement can look low even if both coders captured the same meaning.

One way to address this is to divide the text into fixed units before coding and have coders assign codes only to those units. Question and answer pairs, speaking turns or paragraphs can also serve as units. Whichever route you choose, write the unit definition into the codebook and state it in your report.

Fictional exampleExample (fictional)

A team studying remote work experiences uses the code “workload pressure”. Coder A marks a participant's two-sentence complaint, while Coder B marks a five-sentence passage containing the same complaint. Compared at character level, this looks like a disagreement, yet both coders captured the same meaning. After the pilot round, the team defines the unit as “one speaking turn by the participant” and compares codes on those units.

Measures: from percent agreement to Krippendorff's alpha

Percent agreement is the proportion of units on which coders made the same decision. It is easy to calculate and explain, but it ignores the fact that some agreement would occur by chance even if coders decided at random. With rare codes in particular, percent agreement can be misleadingly high because both coders say “no code” for most units.

Cohen's kappa (Cohen, 1960) addresses this through chance correction. First, observed agreement (po) is calculated. Next, you work out how much agreement would be expected by chance (pe) if each coder's rate of applying the code were independent of the other's. Kappa is the ratio of agreement achieved beyond chance to the maximum agreement possible beyond chance: κ = (po − pe) / (1 − pe). A value of 1 means perfect agreement and 0 means chance-level agreement; negative values indicate agreement below chance. Cohen's kappa was developed for two coders and nominal categories.

Krippendorff's alpha (Krippendorff, 2019) compares observed disagreement with the disagreement expected by chance. Because it works with more than two coders, with missing data (where not every coder codes every unit) and with scale types beyond nominal, it is a common choice in content analysis. Hallgren (2012) offers a tutorial that walks through kappa variants for two or more coders and the intraclass correlation coefficient for continuous ratings, with a focus on choosing the statistic that matches the study design.

Comparison of common measures
MeasureChance correctionNumber of codersStrengthWatch out for
Percent agreementNoUsually twoEasy to calculate and understandCan be misleadingly high for rare codes
Cohen's kappaYesTwoAccounts for chance agreementSensitive to code prevalence and to each coder's rate of applying the code
Krippendorff's alphaYesTwo or moreHandles missing data and different scale typesTedious to compute by hand; needs a suitable statistics tool

Worked example: kappa step by step

The numbers below are entirely fictional and serve only to show the logic of the calculation. Two coders independently decided, for each of 50 transcript segments, whether the code “workload pressure” applies.

Fictional agreement table (50 segments)
Coder A / Coder BB: code presentB: code absentTotal
A: code present18422
A: code absent62228
Total242650
Fictional exampleExample (fictional)

Step 1, observed agreement: the cells where both coders made the same decision are 18 (both “present”) and 22 (both “absent”). po = (18 + 22) / 50 = 40 / 50 = 0.80. Step 2, each coder's rates: A applied the code to 22 of 50 segments (0.44) and not to 28 (0.56). B applied it to 24 (0.48) and not to 26 (0.52). Step 3, agreement expected by chance: both saying “present” is 0.44 × 0.48 = 0.2112; both saying “absent” is 0.56 × 0.52 = 0.2912. pe = 0.2112 + 0.2912 = 0.5024. Step 4, kappa: κ = (0.80 − 0.5024) / (1 − 0.5024) = 0.2976 / 0.4976 ≈ 0.598. Percent agreement is 80%, while chance-corrected agreement is about 0.60.

Interpreting values: benchmarks and their limits

The most widely used classification for interpreting kappa comes from Landis and Koch (1977). The authors offered these ranges as useful benchmarks for discussion while acknowledging that the divisions are arbitrary. McHugh (2012) argues that these labels are too lenient, especially in fields such as health research where decisions have consequences, proposes a stricter interpretation scale and recommends reporting percent agreement alongside kappa.

Two interpretation scales for kappa
ScaleRanges and labels
Landis and Koch (1977)Below 0.00: poor; 0.00–0.20: slight; 0.21–0.40: fair; 0.41–0.60: moderate; 0.61–0.80: substantial; 0.81–1.00: almost perfect
McHugh (2012)0–0.20: none; 0.21–0.39: minimal; 0.40–0.59: weak; 0.60–0.79: moderate; 0.80–0.90: strong; above 0.90: almost perfect

The value of 0.598 in our fictional example falls in the “moderate” range according to Landis and Koch, while on McHugh's scale it sits on the boundary between “weak” and “moderate” depending on rounding. The fact that one number can receive two different labels is a reminder that benchmarks are conventions rather than firm rules. Krippendorff (2019) sets a higher bar for alpha: he treats values around 0.80 as the level at which data can be relied on and values around 0.667 as the lower limit for drawing only tentative conclusions.

Decide which threshold you will use before coding, based on how much coding decisions will affect your conclusions, and justify it. Examining values code by code rather than relying on a single overall coefficient also shows which code definitions need sharpening.

Procedure: from training to revision

The steps below form an adaptable workflow consistent with the recommendations of O'Connor and Joffe (2020) and of MacQueen et al. (1998) on team-based coding. For preparing the codebook itself, see our codebook guide.

  1. Prepare the codebook: For each code, write a name, a definition, inclusion and exclusion rules and short examples; add a definition of the coding unit.
  2. Train the coders: The team reads the codebook together and codes a few segments jointly while thinking aloud; vague definitions surface at this stage.
  3. Run pilot double-coding: Coders independently code a subset of the data, for example segments drawn from different participants and topics. Justify the size of the subset and how it was selected.
  4. Calculate agreement: Calculate your chosen measure code by code, and report percent agreement alongside it.
  5. Hold a disagreement meeting: Discuss every segment coded differently and identify the source of the disagreement: an unclear definition, a unit boundary, an attention slip or a genuine difference in interpretation?
  6. Revise the codebook: Clarify definitions and rules, merge or split codes if needed and write down the rationale for each change.
  7. Run another round if needed: A new subset is coded independently with the revised codebook; once the criterion you set in advance is met, coding is divided among the team.
  8. Document: Record the date of each round, the coders, the unit definition, the calculated values, the meeting decisions and the codebook version.

The number and type of disagreements resolved through discussion are valuable information too: codes that are debated often tend to be either the most conceptually interesting or the most poorly defined.

Reporting and tool support

So that readers can evaluate your coding process, aim to include the following in your methods section:

  • Why you did or did not measure agreement, and how that relates to your approach
  • The number of coders, the training they received and whether they worked independently
  • The unit definition and the size and selection of the double-coded subset
  • The measure used, the calculation tool and code-level or overall values
  • Percent agreement reported together with the chance-corrected value
  • How disagreements were resolved and what changes they led to in the codebook

Summary

  • Agreement is the degree to which coders make the same decisions; reliability is an inference about whether the process is reproducible.
  • Measuring agreement makes sense in codebook-based, team-based studies with large datasets; it does not fit the logic of reflexive thematic analysis.
  • Define the unit before calculating; if unit boundaries differ, agreement can look artificially low.
  • Percent agreement ignores chance; Cohen's kappa suits two coders, and Krippendorff's alpha suits multiple coders and missing data.
  • Benchmark labels are conventions; set and justify your criterion in advance, and report percent agreement and code-level values together.
  • Document pilot double-coding, disagreement meetings and codebook revisions round by round.

References

  1. Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589–597. https://doi.org/10.1080/2159676X.2019.1628806
  2. Campbell, J. L., Quincy, C., Osserman, J., & Pedersen, O. K. (2013). Coding in-depth semistructured interviews: Problems of unitization and intercoder reliability and agreement. Sociological Methods & Research, 42(3), 294–320. https://doi.org/10.1177/0049124113500475
  3. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104
  4. Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34. https://doi.org/10.20982/tqmp.08.1.p023
  5. Krippendorff, K. (2019). Content analysis: An introduction to its methodology (4th ed.). SAGE. https://doi.org/10.4135/9781071878781
  6. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
  7. MacQueen, K. M., McLellan, E., Kay, K., & Milstein, B. (1998). Codebook development for team-based qualitative analysis. Cultural Anthropology Methods, 10(2), 31–36. https://doi.org/10.1177/1525822X980100020301
  8. McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282. https://doi.org/10.11613/BM.2012.031
  9. O'Connor, C., & Joffe, H. (2020). Intercoder reliability in qualitative research: Debates and practical guidelines. International Journal of Qualitative Methods, 19. https://doi.org/10.1177/1609406919899220

Code your first interview today

The free plan carries a pilot study from start to finish. No credit card required.