What is a codebook and why do you need one?
A codebook is a document, updated throughout a study, that lists every code used in the analysis along with its name, definition, application rules and examples. For a researcher working alone, it stands in for memory; in team projects, it is the main tool that lets different people use the same code in the same sense. Miles, Huberman and Saldaña (2020) stress that codes need clear definitions so they can be applied consistently over time and across researchers.
- Consistency: It ensures a code means the same thing in week three as it did on day one.
- Transparency: It shows readers and reviewers which decisions the findings rest on.
- Team communication: It turns disagreements between coders from differences in personal interpretation into discussions about definitions.
- Audit trail: By recording how codes changed, it makes the analysis traceable after the fact.
When should you start the codebook? The best time is when you assign your first codes. The first version can be rough and incomplete; what matters is that definitions are written alongside coding rather than reconstructed from memory once the analysis is over. The codebook records the decisions you make during analysis, not your interpretations; keeping interpretations and open questions in memos makes both documents easier to read.
The components of a code entry
MacQueen et al. (1998) suggest structuring each code entry for team-based analysis into a few fixed fields: the code name, a brief definition, a full definition, guidance on when to use the code, when not to use it, and examples. DeCuir-Gunby et al. (2011) likewise define each code in their interview codebook with a name, a full definition and an example. The table below offers a practical template that combines these fields.
| Field | What to write | Tip |
|---|---|---|
| Code name | A short, distinctive label | Pick a name that cannot be confused with codes at the same level |
| Brief definition | A one-sentence summary | Serves as a quick reminder while coding |
| Full definition | The scope and conceptual meaning of the code | Cite the source if it is based on theory |
| Inclusion criteria | When the code should be applied | Write them in terms that are observable in the text |
| Exclusion criteria | Similar but different situations where the code does not apply | Name the code it is most likely to be confused with |
| Typical example | A quote that clearly represents the code | Choose it from pseudonymized data |
| Atypical example | A borderline quote that still belongs to the code | Shows how contested cases were resolved |
Of these fields, exclusion criteria are the most often neglected. Yet they usually draw the clearest line between two codes. A single sentence such as “Use this code when the participant describes the institution’s response to a technical problem, not the problem itself.” can be more useful than a long full definition.
Follow a consistent rule for code names as well. Writing the parent and child code together, as in “Workload > Double preparation”, preserves context when the code is exported or appears on its own in a table. Avoid names that carry a positive or negative judgment (for example, “Bad management”); descriptive names (for example, “Communication problems with management”) steer coders toward describing a segment rather than judging it. If you use abbreviations, add a list of them at the start of the codebook.
Development approaches: theory, data and hybrid
Where codes come from also shapes how a codebook grows. DeCuir-Gunby et al. (2011) distinguish theory-driven codes from data-driven codes and show that both can sit side by side in the same codebook.
| Approach | Starting point | Strength | Risk |
|---|---|---|---|
| Theory-driven | Existing theory, models or prior studies | Connects findings directly to the literature | Missing data that do not fit the framework |
| Data-driven | Open reading of the transcripts | Foregrounds participants’ own framing | Code proliferation and definitions that settle late |
| Hybrid | A few codes from theory, new ones from data | Balances direction and openness | Codes from the two sources becoming indistinguishable |
In qualitative content analysis, Hsieh and Shannon (2005) call working with initial categories drawn from theory the directed approach, and letting categories emerge from the data the conventional approach. Fereday and Muir-Cochrane (2006) illustrate a hybrid process in which a template codebook prepared in advance is extended with codes that emerge from the data. If you take the hybrid route, keeping each code’s origin in a separate field of the codebook makes reporting easier.
When you start from the data, write the first version after reading and coding a few transcripts; at that stage you can draw on the first-cycle methods in the interview coding guide. When you start from theory, drafting the first version before entering the data and then testing it on a few transcripts shows early on whether the definitions can actually be applied.
Hierarchy depth: how many levels are enough?
Organizing codes into parent codes and child codes keeps related concepts together and makes the code list easier to navigate. But as depth increases, it becomes harder for coders to find the right branch, and similar codes on different branches are more likely to get mixed up.
- Two or three levels are enough for most studies. A deeper tree can be a sign that some child codes are really details of a definition rather than separate codes.
- Decide what the parent code is for. Is it only a heading, or can it also be applied directly to segments that fit none of its children? Write that decision into the codebook.
- Keep sibling codes at the same level of abstraction. A topic code and an emotion code should not sit side by side under the same parent.
- Do not confuse the hierarchy with themes. The hierarchy organizes codes; themes are interpretive patterns that can cut across several branches.
Saldaña (2025) describes how some researchers work with large chunks of data and few codes, while others work with fine-grained segments and many codes. Whichever way you lean, a tidy hierarchy is one way to keep the consequences of that preference manageable.
Iterative refinement, version control and audit trail
A codebook is not finished with its first version. Guest, MacQueen and Namey (2012) treat the codebook in applied thematic analysis as a document that is revisited as coding progresses. Recording what each change was and why it was made keeps the analysis traceable later on.
| Operation | When | What to record |
|---|---|---|
| Merge | Two codes cannot be told apart in practice | The merged codes, the new definition and the rationale |
| Split | One code carries more than one distinct meaning | Definitions of the new codes and how existing segments were redistributed |
| Rename | The name no longer reflects the definition | Old and new name; whether the definition changed |
| Move | The code sits under the wrong parent | Old and new location |
| Archive | The code is no longer used but its history should be kept | Why it was dropped and what happened to earlier codings |
- Version number and date: Increase the codebook version after every meaningful change (for example 1.0, 1.1, 2.0).
- Change log: Keep the date, the person who made the change, the operation and the rationale on one short line.
- Retroactive coding: If a definition changed, review previously coded segments against the new definition and note that in the log.
- Archive instead of deleting: Deleting an unused code also deletes the record of why it existed.
A line from a change log: Version 1.2. The code “Technical problem” was split into “Infrastructure > Connection problem” and “Infrastructure > Inadequate device”. Rationale: in pilot coding, two coders had applied this code to very different segments; some described the school’s infrastructure, others the student’s device. All segments under the old code were reviewed against the new definitions.
Using a codebook in a team
When a team codes together, the codebook is both training material and the common ground on which disagreements are resolved. DeCuir-Gunby et al. (2011) illustrate a process in which the codebook develops as team members code together, discuss the results and revise definitions.
- Introduction: Go through the codebook field by field with the coders, discussing the exclusion criteria in particular with examples.
- Pilot coding: Ask everyone to code the same few transcripts independently.
- Comparison: Talk through the differences one by one and note the source of each disagreement (unclear definition, different unit or simple oversight).
- Revision: Fix unclear definitions, merge or split codes if needed and share the new version.
- Regular checks: While coding continues, compare coding on shared segments again at set intervals.
Whether to measure agreement numerically is a methodological choice. O'Connor and Joffe (2020) summarize the debates about the place of intercoder reliability in qualitative research and offer practical guidelines for applying it when it is used. For a detailed treatment, see the inter-coder agreement guide.
Sample codebook
The excerpt below was prepared for a fictional study of middle school teachers’ experiences with hybrid teaching. The quotes are invented and serve only to show how code entries can be written.
| Code | Brief definition | Inclusion criteria | Exclusion criteria | Typical example |
|---|---|---|---|---|
| Workload > Double preparation | Preparing separately for classroom and online teaching | Describes preparing two separate materials or plans for the same lesson | General fatigue or long working hours (use “Workload > After-hours work”) | “I prepared every topic twice, once for the board and once for the screen.” |
| Workload > After-hours work | Work that spills beyond teaching hours | Describes work done in the evening, on weekends or during holidays | Heavier workload within working hours | “I was answering parents’ messages at midnight.” |
| Student engagement > Loss of contact | Weakening interaction with students | Not being able to see students, get their reactions or know who one is talking to | A student unable to join for technical reasons (use “Infrastructure > Connection problem”) | “With the cameras off, it felt like I was talking into a void.” |
| Institutional support > Technical-only training | Training provided by the institution is limited to using tools | States that the training covered software, platform or device use | General statements about a lack of pedagogical support | “They showed us the buttons in the program, and that was it.” |
| Peer support | Giving or receiving informal help from colleagues | Teachers sharing information, materials or encouragement among themselves | Formal meetings organized by school management | “We set up a message group in our department and sorted everything out there.” |
The full definition and atypical example fields, which do not fit in the table, are kept in a separate entry for each code in the extended codebook. For “Student engagement > Loss of contact”, for example, an atypical example might be: “The cameras were on, but nobody was looking at me.” Although the cameras were on, the teacher describes being unable to make contact, so the segment is included in the code.
Common mistakes
Problems in a codebook usually surface only once coding is under way and disagreements start to pile up. You can use the list below as a checklist after writing the first version and at every major revision.
- Writing only code names: Undefined codes become empty boxes that each coder fills with their own interpretation.
- Skipping exclusion criteria: If the boundary between similar codes is not written down, the same segment ends up under different codes for different coders.
- Confusing codes with themes: Writing interpretive claims into the codebook can lead themes to close too early, without being questioned.
- Overly deep hierarchies: Trees with four or five levels slow coding down and increase inconsistency.
- Not recording changes: If the rationale for merges and splits is not written down, the analysis cannot be explained later.
- Changing a definition but leaving earlier coding as is: The same code ends up meaning different things in different parts of the data.
- Writing the codebook after coding is finished: A codebook written after the fact documents only the outcome, not the process.
- Leaving real names in examples: A codebook is often shared, so example quotes should be pseudonymized too.
The codebook in YouReply Qualitative
In YouReply Qualitative, the codebook is hierarchical, and each code can carry a definition along with inclusion and exclusion rules. You can merge, split, move, rename and archive codes; the last operation can be undone, merge and split operations can carry a rationale note, and the operation history is kept. The codebook can be exported to Word and Excel.
Workspaces have admin, researcher, coder and viewer roles. The tool does not compute an inter-coder agreement coefficient, so if you plan an agreement check, you will need to organize it separately. Codebook operations are described in more detail on the analysis workflow page.
Summary
- Give every code entry a name, brief and full definitions, inclusion and exclusion criteria, and typical and atypical examples.
- Record in the codebook whether each code comes from theory or from the data.
- Keep the hierarchy to two or three levels in most cases, and do not confuse codes with themes.
- Log merges, splits, renames and archiving decisions with a version number and rationale.
- In teams, develop the codebook together through pilot coding and regular comparison.
References
- DeCuir-Gunby, J. T., Marshall, P. L., & McCulloch, A. W. (2011). Developing and using a codebook for the analysis of interview data: An example from a professional development research project. Field Methods, 23(2), 136–155. https://doi.org/10.1177/1525822X10388468
- Fereday, J., & Muir-Cochrane, E. (2006). Demonstrating rigor using thematic analysis: A hybrid approach of inductive and deductive coding and theme development. International Journal of Qualitative Methods, 5(1), 80–92. https://doi.org/10.1177/160940690600500107
- Guest, G., MacQueen, K. M., & Namey, E. E. (2012). Applied thematic analysis. SAGE. https://doi.org/10.4135/9781483384436
- Hsieh, H.-F., & Shannon, S. E. (2005). Three approaches to qualitative content analysis. Qualitative Health Research, 15(9), 1277–1288. https://doi.org/10.1177/1049732305276687
- MacQueen, K. M., McLellan, E., Kay, K., & Milstein, B. (1998). Codebook development for team-based qualitative analysis. Cultural Anthropology Methods, 10(2), 31–36. https://doi.org/10.1177/1525822X980100020301
- Miles, M. B., Huberman, A. M., & Saldaña, J. (2020). Qualitative data analysis: A methods sourcebook (4th ed.). SAGE.
- O'Connor, C., & Joffe, H. (2020). Intercoder reliability in qualitative research: Debates and practical guidelines. International Journal of Qualitative Methods, 19. https://doi.org/10.1177/1609406919899220
- Saldaña, J. (2025). The coding manual for qualitative researchers (5th ed.). SAGE. https://doi.org/10.4135/9781036235611