Skip to main content

What is qualitative data analysis? Methods and steps

Qualitative data analysis is the systematic work of making sense of interviews, focus groups, open-ended answers and documents. This guide brings together the main approaches, the stages of analysis, trustworthiness criteria and the reporting standards reviewers expect.

Prepared by:
YouReply Qualitative content team
Published:
Last updated:
{minutes} min read
9 min read

What is qualitative data analysis?

Qualitative data analysis is the process of systematically reading and organizing non-numeric data (transcripts, written responses, field notes, documents) to develop patterns of meaning that answer a research question. The goal is less about measuring how often something occurs and more about understanding how people make sense of an experience, and what they do, in which context, and why.

The key difference from quantitative analysis is that the work is largely interpretive. Researchers do not simply sort data into boxes; through codes, categories and themes they bring a reading to the data. That is why the researcher's position, the decisions they make and the reasoning behind those decisions matter as much as the findings themselves.

The process is usually cyclical rather than linear. First notes are taken during data collection, code definitions shift during coding, and developing themes sends you back to the data again. Different qualitative traditions, such as narrative research, phenomenology, grounded theory, ethnography and case study, frame analysis in different ways; your research question and design should drive which approach you choose.

What kinds of data can you analyze qualitatively?

Qualitative analysis works with any data where meaning is carried in words, stories and context. The type of data also shapes what you need to watch for during analysis.

Common qualitative data types and what to watch for
Data typeTypical sourceWhat to watch for in analysis
Individual interviewSemi-structured interview transcriptAlso read how the interviewer's questions shaped the answers.
Focus groupMulti-speaker transcriptSeparate speakers accurately; the interaction within the group is data too.
Open-ended survey responseShort written answersAnswers can be brief and stripped of context, which limits deep interpretation.
Field notes and observationResearcher notesKeep description and your own interpretation apart.
DocumentMinutes, policy texts, correspondenceConsider who produced the document and for what purpose.

If you work with interviews, remember that a transcript is itself an interpretation: how you render pauses, laughter or overlapping speech affects the analysis. Setting and writing down your transcription conventions before you start makes later consistency much easier. For the details of coding interview data, see the interview coding guide.

The main analytic approaches

Qualitative analysis is not a single technique but a family of methods that use similar tools (codes, categories, themes, analytic memos) for different purposes. Four approaches come up most often.

Thematic analysis. Braun and Clarke (2006) describe thematic analysis as a flexible method for identifying, analyzing and reporting patterns of meaning (themes) across a dataset. Because it is not tied to a particular theoretical framework, it is used widely in health, education, psychology and user research. For a step-by-step walkthrough, read how to do thematic analysis.

Qualitative content analysis. Hsieh and Shannon (2005) distinguish three approaches. In conventional content analysis, categories are derived directly from the data. In directed content analysis, an existing theory or prior research determines the initial codes. In summative content analysis, specific words or content are counted first and their use in context is then interpreted.

Grounded theory. In Charmaz's (2014) constructivist grounded theory, analysis runs alongside data collection. Initial coding, focused coding, constant comparison and memo writing move the work toward a conceptual account that explains a process. The aim is not only to describe but to build an explanation grounded in the data.

Matrix-based analysis. Miles, Huberman and Saldaña (2020) link data condensation, data display, and drawing and verifying conclusions. Matrices with participants or cases in the rows and codes or categories in the columns make it easier to compare across cases and groups.

A quick comparison of approaches
ApproachCore questionTypical outputWhen it fits
Thematic analysisWhat patterns of shared meaning are in the data?Defined themes and an analytic narrativeQuestions about experience, perception and sense-making
Qualitative content analysisWhich categories appear, and in what context?A category system, with frequencies where usefulSystematic description of text, testing existing theory
Grounded theoryHow does this process work?A conceptual model or theoryUnderstudied processes, theory development
Matrix-based analysisWhat differs across cases and groups?Case × code matrices, comparative summariesApplied, comparative and team-based projects

The stages of analysis

Whatever approach you choose, most qualitative analysis follows a similar skeleton. The sequence below is not a rigid checklist but a road map that expects you to loop back.

  1. Preparation: Finish your transcripts, remove identifying details or assign pseudonyms, and keep participant attributes (age group, role, organization) in a separate table.
  2. Familiarization: Read all the data from start to finish at least once and note first impressions; do not rush into coding at this point.
  3. Coding: Give short, meaningful labels to the pieces of data that bear on your research question. The codebook guide shows how to write code definitions and inclusion and exclusion rules.
  4. Developing categories and themes: Group codes by how they relate; ask whether a group carries shared meaning rather than just a recurring topic.
  5. Reviewing: Check themes against both the coded excerpts and the dataset as a whole; look deliberately for contradictory and outlying cases.
  6. Interpreting and writing: Relate findings to your research question and the literature; use quotes to support your argument, not only to illustrate it.
Fictional exampleExample (fictional)

Imagine a fictional interview study about a city's digital services. On first reading you notice that many participants talk about “waiting in line.” During coding, that expression splits into codes such as “waiting time,” “uncertainty” and “loss of control.” While developing themes, you see that the real pattern is not waiting itself but not knowing where an application stands, and you name the theme accordingly. Back in the dataset, you find a participant who waited a long time yet was satisfied and who describes being kept informed; rather than disproving your theme, this outlier sharpens its boundaries.

Keeping analytic memos throughout the stages grounds the analysis in records rather than memory. Writing a few sentences at the moment you create a code, merge two codes or drop a theme keeps later decisions consistent and leaves you a ready-made audit trail when you write the methods section. In a team, these memos also give coders common ground for discussing how each of them understands the same code.

Rigor and trustworthiness

Quality in qualitative research is not judged by transplanting the quantitative notions of validity and reliability one to one. Lincoln and Guba (1985) proposed credibility, transferability, dependability and confirmability instead. Nowell and colleagues (2017) adapt these criteria to the phases of thematic analysis and illustrate how trustworthiness can be demonstrated at each step.

Trustworthiness criteria and what they look like in practice
CriterionQuestion it asksExample in practice
CredibilityDo the findings plausibly reflect participants' experience?Prolonged engagement with the data, triangulation, peer debriefing, negative case analysis
TransferabilityCan readers judge whether the findings fit other contexts?Thick description of participants and setting
DependabilityIs the process traceable and logical?Codebook versions, dated decision memos, an audit trail
ConfirmabilityAre interpretations grounded in the data?Excerpts linked to source text, reflexive memos

In day-to-day practice, three habits strengthen trustworthiness: recording important decisions such as merging codes or renaming a theme in a dated memo, making your own assumptions visible through reflexive memos, and reporting cases that contradict your themes instead of hiding them. Sharing findings with participants is another option, but since it is not always realistic to expect participants to endorse analytic interpretation that goes beyond their own experience, decide up front what that step is for.

Reporting standards: COREQ and SRQR

Two checklists are widely used to report qualitative studies transparently. COREQ, developed by Tong, Sainsbury and Craig (2007), is a 32-item checklist for interview and focus group studies, grouped into three domains: research team and reflexivity, study design, and analysis and findings. SRQR, developed by O'Brien and colleagues (2014), is a 21-item standard synthesized from existing recommendations and designed to apply across qualitative approaches.

Both rest on the same principle: readers should be able to judge how the findings were produced. At a minimum, your report should make the following clear:

  • Who the researchers are, their relationship to participants and their perspective
  • The sampling strategy, number of participants and data collection context
  • How transcription, coding and theme development were carried out
  • How many people coded and how differences were handled
  • Findings supported by quotes, with each quote tagged by participant code

Read whichever checklist your journal or institution requires before you begin analysis. Much of what you need to report (decision memos, codebook versions, participant attributes) will only exist later if you keep records throughout the process.

Common mistakes

Most of the problems below come from losing sight of the interpretive nature of analysis, or of the need to make methodological decisions visible.

  • Presenting interview questions as themes: Reporting your question headings as themes is not analysis; it is reordering the data.
  • Writing that themes “emerged”: Themes are developed through the researcher's analytic decisions, and this language hides the process.
  • Equating frequency with importance: A code that appears often is not automatically important; what it says about the research question is what counts.
  • Leaving the codebook undocumented: Codes without definitions and rules drift over time and get read differently across a team.
  • Leaving out contradictory cases: Examples that do not fit a theme are often the most valuable evidence of where its boundaries lie.
  • Method and practice diverging: Naming one approach in the methods section while following another approach's procedures makes it hard for readers to evaluate the findings.

Most of these mistakes are prevented at the start of analysis rather than the end: by sharpening the research question, choosing an approach deliberately and keeping decisions in writing along the way.

How software helps, and where it stops

Qualitative analysis software does not do the analysis for you; it carries the structure of the analysis. Coding hundreds of excerpts, merging codes, tracking which participants a theme draws on, or finding the reasoning behind a decision months later quickly becomes unmanageable with paper and scattered spreadsheets. Software helps most with:

  • Keeping coded excerpts linked to the source text so you can retrieve them with their context
  • Applying codebook changes such as merging, splitting and renaming consistently
  • Quickly producing matrices that compare codes or themes with participant attributes
  • Linking memos to the relevant code, theme or excerpt so an audit trail builds up

The limits matter just as much. Software cannot decide whether a code is apt or whether a theme truly answers the research question. AI-assisted suggestions can speed up drafting, but they can miss context, present wording that does not quite appear in the data as if it were a quote, or pick up surface topics more readily than interpretive themes. Every suggestion should therefore be checked against the source text, and the final call should stay with the researcher. For the data protection and transparency side, see ethics in AI-assisted qualitative analysis.

How YouReply Qualitative supports this process

YouReply Qualitative is designed to run these stages in a single browser-based workspace. You can upload .docx, .pdf files with a text layer, .txt, .md, .vtt, .srt and .csv files, paste text, or import open-ended answers from YouReply Survey. You do the coding yourself by highlighting text; the codebook holds definitions along with inclusion and exclusion rules, and themes sit as a separate layer above codes.

AI suggestions are optional per study and never enter counts, matrices or reports unless you accept them. Analytic, methodological and reflexive memos can be attached to a code, theme or excerpt, and the codebook, quote table, matrices and report export to Word and Excel. See the analysis workflow for details.

Summary

  • Qualitative data analysis is a systematic, interpretive way to develop patterns of meaning from text that answer a research question.
  • Thematic analysis, qualitative content analysis, grounded theory and matrix-based analysis answer different questions; let your research question choose.
  • Preparation, familiarization, coding, theme development, reviewing and writing are cyclical rather than linear.
  • Trustworthiness is shown through credibility, transferability, dependability and confirmability, backed by a well-kept audit trail.
  • COREQ and SRQR remind you what to report transparently; record that information as you go.
  • Software provides structure and traceability; interpretation and the final decision stay with the researcher.

References

  1. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
  2. Charmaz, K. (2014). Constructing grounded theory (2nd ed.). SAGE.
  3. Hsieh, H.-F., & Shannon, S. E. (2005). Three approaches to qualitative content analysis. Qualitative Health Research, 15(9), 1277–1288. https://doi.org/10.1177/1049732305276687
  4. Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. SAGE.
  5. Miles, M. B., Huberman, A. M., & Saldaña, J. (2020). Qualitative data analysis: A methods sourcebook (4th ed.). SAGE.
  6. Nowell, L. S., Norris, J. M., White, D. E., & Moules, N. J. (2017). Thematic analysis: Striving to meet the trustworthiness criteria. International Journal of Qualitative Methods, 16(1). https://doi.org/10.1177/1609406917733847
  7. O'Brien, B. C., Harris, I. B., Beckman, T. J., Reed, D. A., & Cook, D. A. (2014). Standards for reporting qualitative research: A synthesis of recommendations. Academic Medicine, 89(9), 1245–1251. https://doi.org/10.1097/ACM.0000000000000388
  8. Tong, A., Sainsbury, P., & Craig, J. (2007). Consolidated criteria for reporting qualitative research (COREQ): A 32-item checklist for interviews and focus groups. International Journal for Quality in Health Care, 19(6), 349–357. https://doi.org/10.1093/intqhc/mzm042

Code your first interview today

The free plan carries a pilot study from start to finish. No credit card required.