Before coding: preparing the transcript
Coding happens on the transcript, so the way a transcript is prepared shapes what you can and cannot see during analysis. Brinkmann and Kvale (2015) emphasize that transcription translates spoken conversation into written text and inevitably involves choices: there is no single “correct” transcript, only one that fits the research question. If you study discourse or interaction, pauses and overlapping talk matter; for a content-focused thematic study, a readable, lightly cleaned text is often enough.
Whatever level of detail you choose, write the conventions down at the start of the study and apply them the same way across all transcripts. When more than one person transcribes, a short transcription guide greatly reduces inconsistency.
| Situation | Suggested notation | Note |
|---|---|---|
| Change of speaker | “Interviewer:” or “P3:” at the start of the line | Each turn starts on a new line |
| Inaudible passage | [inaudible 00:14:32] | A timestamp makes it easy to go back to the recording |
| Long pause | [pause] | No need to record duration unless you analyze interaction |
| Nonverbal reaction | [laughs], [sighs] | Include it when it changes the meaning |
| Overlapping talk | [interrupts] | Especially important in focus groups |
| Identifying information | [employer], [town] | Generalized in square brackets during pseudonymization |
- Personal names: Refer to participants with codes such as P1 and P2 or with consistent pseudonyms; also replace the names of third parties mentioned in the text.
- Places and organizations: Use generalizations such as “[a private hospital]” or “[a small coastal town]”.
- Indirect identifiers: A rare occupation, job title, event or combination of dates can reveal identity on its own; generalize it if it adds nothing to the analysis.
- Key file: Store the mapping between real names and codes separately from the transcripts, in a place with restricted access.
Coding units and the unitization problem
The coding unit is the piece of text to which you assign a code or set of codes. In open-ended survey answers, the unit is usually the answer itself; in semistructured interviews, however, the text does not break naturally into pieces. A participant may touch on three different topics in a single turn or stretch one idea across five minutes.
Campbell et al. (2013) discuss this as the problem of unitization: even when two coders apply the same code to the same passage, choosing different start and end points makes agreement hard to assess. The authors propose a procedure that standardizes the units of text coders work on and then refines the coding scheme until it reaches an acceptable level of reliability or agreement. If you work as a team, it is therefore important to put your unit decision in writing; see the inter-coder agreement guide for details.
| Unit | When it fits | Watch out for |
|---|---|---|
| Line or sentence | Detailed, inductive first reading | Context is easily lost; the number of codes grows fast |
| Meaning unit | Most thematic and content-focused studies | Boundaries are open to interpretation; write a rule |
| Speaker turn | Focus groups, questions about interaction | Long turns can mix several ideas |
| Question-answer block | Structured interviews, comparative summaries | Stories that go beyond the question may be missed |
A practical rule: keep a segment wide enough to make sense when lifted out of context, but narrow enough to carry a single idea. Braun and Clarke (2006) recommend coding extracts inclusively, keeping some of the surrounding text, so that you do not have to reconstruct the context when you return to the segment later.
Deductive, inductive and hybrid approaches
Where your codes come from depends on your research question and theoretical framework. In a deductive approach, codes are derived from theory, prior studies or the topics in your interview guide; Miles, Huberman and Saldaña (2020) describe how such a provisional start list can give direction to analysis. In an inductive approach, codes emerge while you read the data. Hsieh and Shannon (2005) distinguish these two directions in qualitative content analysis as the directed and the conventional approach, respectively.
- Start deductively if you are testing a specific model or extending earlier findings.
- Start inductively if the topic is understudied and you want to foreground participants’ own framing.
- Choose a hybrid approach if you want to begin with a few core codes from theory while staying open to new codes that appear in the data.
Fereday and Muir-Cochrane (2006) illustrate, step by step, a hybrid process that uses a theory-derived template codebook together with codes that emerge from the data. Whichever route you take, writing definitions for your starting codes keeps their meaning from drifting over time; the codebook guide explains how to do that in detail.
First-cycle and second-cycle coding
Saldaña (2025) thinks of coding in cycles. In the first cycle, the data receive their initial labels; in the second cycle, those labels are reorganized, grouped and developed into more abstract categories. The author describes many methods; the table below gives a general overview of a few that are often used with interview data.
| Method | Cycle | What it captures | Fictional code example |
|---|---|---|---|
| Descriptive coding | First | The topic of a segment in a short word or noun phrase | Screen time |
| In vivo coding | First | A participant’s own striking phrase, in quotation marks | “The workday never ends” |
| Process coding | First | Action and change, usually with gerunds (-ing words) | Taking work home |
| Values coding | First | The participant’s values, attitudes and beliefs | Valuing face-to-face contact |
| Pattern coding | Second | Gathers first-cycle codes under an explanatory meta-code | Blurring boundaries |
These methods are not mutually exclusive; you can apply descriptive and in vivo codes to the same transcript. Charmaz (2014) describes a similar movement in grounded theory when she moves from line-by-line initial coding to focused coding: first many labels that stay close to the data, then a selection of the most significant and frequent among them.
The goal of the second cycle is not only to reduce the number of codes but to see how codes relate to each other. A pattern code proposes an explanation for why several first-cycle codes appear together; such explanations may later develop into themes. For the move from codes to themes, see the thematic analysis guide.
Writing memos while coding and overlapping codes
Memos are short notes in which you record the interpretations, questions and decisions that come up during analysis. Both Saldaña (2025) and Charmaz (2014) treat memo writing not as a task separate from coding but as an integral part of it: a code names a segment, and a memo explains why that name was chosen and what it points to.
- Analytic memo: “P2 and P5 describe technical problems as personal failures; this may be a pattern.”
- Methodological memo: “Question-answer blocks were too long; I switched the unit to meaning units.”
- Reflexive memo: “My own teaching experience may be leading me to read this segment negatively.”
Overlapping codes means applying more than one code to the same piece of text, and this is entirely normal in interview data: one sentence can carry both a topic and a value. Braun and Clarke (2006) also note that an extract can be coded into as many codes or themes as it fits. The problem is not the overlap itself but ambiguity.
- If you give a segment two codes, make sure the two codes really capture different things.
- If two codes almost always appear together, consider merging them or sharpening their definitions.
- If one code is a subtype of another, build a hierarchy instead of overlapping them.
- If one code usually covers a wider passage (for example, a whole story and a single sentence within it), write that relationship into the code definitions.
Coding focus group recordings
Focus group data are an interaction in which several people respond to one another, so they call for different decisions than individual interviews. The first decision is the unit of analysis: are you analyzing individuals, the group, or both?
- Speaker identity: Every turn should be linked to the right participant. If voices are hard to tell apart in the recording, mark uncertain turns in the transcript.
- Code the interaction: Moves such as agreeing, objecting or building on the previous speaker can be as meaningful as the content.
- Dominant voices: One or two talkative people can make a code look “common in the group”; check how many different participants contributed to it.
- The illusion of consensus: Do not assume that silent participants agree; note silence separately.
- Participant attributes: If you keep information such as age, role or region at the participant level, you can compare codes across those attributes.
Worked example: coding a short interview excerpt
Imagine a fictional study of middle school teachers’ experiences with hybrid teaching. Here is an excerpt from the interview with participant P4: Interviewer: How did your day change when you moved to hybrid teaching? P4: The first month was chaos. In the evenings I prepared two separate lesson plans, one for the classroom and one for the screen. [pause] It felt like the workday never ends. But what wore me out most was teaching students with their cameras off; I didn’t know who I was talking to. Interviewer: Did you get support from the school? P4: They gave us a training, but it was about how to open the system, not how to build a lesson. In the end my colleagues and I set up our own group.
For this excerpt, let us use the meaning unit as the coding unit and take a hybrid approach: the starting code “Institutional support” comes from the interview guide, and the other codes emerge from the data.
| Segment | Code | Method |
|---|---|---|
| “In the evenings I prepared two separate lesson plans” | Taking work home; Double preparation load | Process; descriptive |
| “It felt like the workday never ends” | “The workday never ends” | In vivo |
| “teaching students with their cameras off; I didn’t know who I was talking to” | Losing connection with students; Valuing face-to-face contact | Descriptive; values |
| “about how to open the system, not how to build a lesson” | Institutional support: technical-only training | Deductive starting code |
| “my colleagues and I set up our own group” | Seeking peer support | Process |
While coding, you might attach an analytic memo like this to the segment: “P4 makes up for missing pedagogical support through colleagues. Do other interviews show the same substitution of peer support for formal support?” In the second cycle, if “Taking work home”, “Double preparation load” and “The workday never ends” also appear together in other interviews, they could be gathered under a pattern code such as “Blurring boundaries”. This is not yet a theme; it is a proposed explanation that needs to be tested across several participants.
Common mistakes
Each of the mistakes below can look small on its own, but together they weaken how trustworthy and reportable the analysis is. Reviewing this list during the first week of coding can spare you major corrections later.
- Undefined codes: The meaning of a code that has only a name drifts within weeks. Write at least a one-sentence definition for every code.
- Code proliferation: Opening a new code for every segment produces hundreds of codes used only once. Before creating a code, check whether an existing one fits.
- Cutting segments off from their context: Very short segments can take on different meanings when read later.
- Coding only the interesting parts: Skipping passages that contradict your expectations or look ordinary makes the analysis selective.
- Mistaking a code for a theme: “Technical problems” summarizes a topic; a theme makes a claim about a meaningful pattern in the data.
- Equating frequency with importance: A code applied many times is not necessarily the most important finding for your research question.
- Changing a definition without revisiting earlier coding: When a code’s definition changes, review the segments already coded with it.
- Leaving pseudonymization to the end: Once real names have spread into quotes and memos, removing them is both hard and risky.
Coding interviews with YouReply Qualitative
YouReply Qualitative is a browser-based qualitative analysis workspace with Turkish and English interfaces. You can upload transcripts as .docx, .pdf with a text layer, .txt, .md, .vtt, .srt or .csv (up to 20 MB per file) or paste text; there is no audio or video transcription and no OCR, so the transcript needs to be prepared beforehand. When the transcript uses “Name:” style speaker labels, turns are split automatically; in focus groups you can map speakers to participants and give each participant a code such as P1 and an optional pseudonym.
Coding is done manually by highlighting text, and a segment can carry more than one code; memos can be attached directly to an excerpt, a code or a participant. AI suggestions, which can be switched on per study, stay pending until you accept, edit or reject them, and unaccepted suggestions never enter counts. Because speaker labels are sent to the AI provider as written when AI is on, replacing real names before upload is recommended. See the analysis workflow page for details.
Summary
- Put transcription conventions and pseudonymization in writing before you start coding.
- Choose the coding unit deliberately and, in team projects, record that decision in the codebook.
- Use labels close to the data in the first cycle and gather them into pattern codes in the second.
- Write memos while coding: the code names a segment, the memo explains why.
- In focus groups, code the interaction too, and do not confuse the number of codings with the number of people.
References
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
- Brinkmann, S., & Kvale, S. (2015). InterViews: Learning the craft of qualitative research interviewing (3rd ed.). SAGE.
- Campbell, J. L., Quincy, C., Osserman, J., & Pedersen, O. K. (2013). Coding in-depth semistructured interviews: Problems of unitization and intercoder reliability and agreement. Sociological Methods & Research, 42(3), 294–320. https://doi.org/10.1177/0049124113500475
- Charmaz, K. (2014). Constructing grounded theory (2nd ed.). SAGE.
- Fereday, J., & Muir-Cochrane, E. (2006). Demonstrating rigor using thematic analysis: A hybrid approach of inductive and deductive coding and theme development. International Journal of Qualitative Methods, 5(1), 80–92. https://doi.org/10.1177/160940690600500107
- Hsieh, H.-F., & Shannon, S. E. (2005). Three approaches to qualitative content analysis. Qualitative Health Research, 15(9), 1277–1288. https://doi.org/10.1177/1049732305276687
- Miles, M. B., Huberman, A. M., & Saldaña, J. (2020). Qualitative data analysis: A methods sourcebook (4th ed.). SAGE.
- Saldaña, J. (2025). The coding manual for qualitative researchers (5th ed.). SAGE. https://doi.org/10.4135/9781036235611