A codeframe is the list of categories a researcher uses to turn free-text survey answers into something you can count. Ask a thousand people "what could we do better?" and you get a thousand sentences. Apply a codeframe and you get a table: 34% mentioned price, 19% mentioned delivery speed, 11% mentioned a specific feature. The codeframe is what makes that second thing possible.
It sounds like a small piece of infrastructure, and it's one of the few things in survey research where a bad decision at the start is expensive to fix later — every wave, every report, every cross-tab built on top of it inherits whatever the codeframe got wrong.
What a codeframe actually is
Structurally, a codeframe is a flat or hierarchical list: a set of named categories, each with a short definition of what belongs in it, applied to one open-ended question. "Price," "Customer service," "Product quality," "Packaging," "Other" is a codeframe for a satisfaction question. Each category on the list is a code; coding a response means assigning it one or more of those codes.
The term comes from content analysis and shows up under slightly different names depending on who you ask: coding frame in UK usage, code list or category scheme in some academic writing, and just "the codebook" in a lot of commercial research where the two words are used as synonyms. That overlap is worth untangling, because in careful usage they aren't quite the same thing.
Codeframe vs. codebook
In the strict sense, the codeframe is the list of categories — the names and the hierarchy. The codebook is the codeframe plus everything a coder needs to apply it consistently: a definition for each code, a handful of example responses, and the edge-case rules that got decided the second week of coding ("mentions of a store employee by name go under Customer Service, not the employee's own catch-all"). Most working researchers say "codebook" for the whole document and "codeframe" for the category structure inside it, and use the words interchangeably the rest of the time. Nobody will correct you for either.
What actually matters isn't the label but whether the artefact your coders — human or AI — are working from has that second part: definitions and examples, not just a list of names. A list of names is a guess at consistency. A definition with two or three real examples is what makes two different coders, or the same coder on a different day, land on the same code for the same response.
The anatomy of a codeframe
A codeframe for "What could we do to improve?" on a retail satisfaction survey might look like this:
Code Definition
Price Mentions of cost, value for money, or wanting discounts
Product quality Complaints or praise about durability or how the product performs
Delivery / shipping Speed, cost, damage, or tracking of delivery
Customer service Interactions with staff, support, or the returns process
Website / app Usability of the online store or app
Other Doesn't fit any of the above; too vague to code
Two things about that list are doing more work than they look like they are. The Other row is a code like any other, and it needs a base rate you check: if 30% of responses end up there, the codeframe is missing a real category — it isn't describing an unusually vague sample. And the definitions are short on purpose. A definition that takes a paragraph to read is a definition nobody applies consistently under time pressure.
Real codeframes are rarely this short. A well-run open-end on a large tracking study can run twenty to sixty codes, and at that size a flat list stops being usable — which is why most coding platforms, ours included, support grouping codes under a category and a subcategory (Product quality › Durability, Product quality › Performance) instead of forcing sixty items into one undifferentiated list. The hierarchy doesn't change what gets counted; it changes whether a person can scan the codeframe and find the code they need in ten seconds instead of two minutes.
What makes a codeframe good
Four properties, in the order they tend to get discovered as mistakes rather than planned as goals:
- The categories don't overlap. If a response about "the app crashed during checkout" could plausibly go under either Website/App or Customer Service, one coder will pick one and another coder will pick the other — and the split between those two codes will end up measuring coder behavior, not respondent behavior.
- Together they cover what people actually said, not what you expected them to say. A codeframe drafted from the questionnaire designer's intuition, before reading a single response, reliably misses the two or three themes that turn out to matter most. The fix is mechanical: read fifty to a hundred real responses before finalizing the list.
- Each code has a real base rate. A code that catches two responses out of two thousand isn't a category, it's a note — worth merging into a broader code or leaving in "Other" rather than giving it a row in every future cross-tab.
- Definitions exist and examples back them up. Not for documentation's sake — for the moment, eight weeks from now, when someone asks why a particular response is coded the way it is, and the answer needs to be a rule, not a shrug.
How a codeframe gets built
Three starting points, and they aren't mutually exclusive on a real project:
Built by hand from a sample. A researcher reads a batch of responses — usually a hundred to a few hundred — and drafts categories inductively, the way qualitative coding has always worked. Thorough, slow, and still the right call for a study with no precedent to build on.
Generated from the actual data. Rather than starting from a generic template, a model reads a sample of the real open ends for that specific question and proposes a codeframe from what respondents actually wrote — categories, definitions and starting examples included, ready for a researcher to edit before coding starts. This is faster than the manual version and, because it's reading your responses rather than a template, tends to catch a study-specific theme a generic list would have missed. It's also not a substitute for someone reviewing the result: an AI-proposed codeframe is a strong first draft, not a finished one.
Imported from an existing document. Agencies and in-house teams often already have a codeframe — a codebook workbook from a previous wave, or a codeframe a client hands over as part of the brief. Importing one means every code and definition already exists; the work is mapping columns (Category, Code, Definition, Examples) rather than writing them from scratch, and on a tracking study the same workbook gets reused wave over wave so the codeframe doesn't quietly drift.
What happens when a response doesn't fit
No codeframe survives contact with the full dataset unchanged. Somewhere past response two hundred, something gets said that doesn't fit any code on the list. Three things can happen to it, and which one is right depends on the project:
- It goes to "Other." Correct when it's a true one-off — vague, off-topic, or too rare to justify a new category.
- The codeframe grows a new code. Correct when the same theme shows up more than a handful of times — the codeframe was incomplete, not the response weird. This is the case worth catching, because it's easy to miss: nobody notices a theme is under-covered by reading one response, only by noticing that "Other" is unusually large or that a review pass keeps flagging the same kind of comment.
- Nothing changes, on purpose. On a locked tracking codebook — the one thing every wave has to share for trend lines to mean anything — adding a code mid-study is sometimes the wrong call even when the theme is real, because it breaks comparability with every prior wave. The right move there is to log it for the next redesign, not the current one.
A coding platform's job at this stage is to make the first two visible rather than silent: flag what doesn't fit an existing code, let a researcher decide case by case whether it's noise or a gap, and — this is the part that's easy to get wrong — apply that decision retroactively to every earlier response with the same theme, not just the one that triggered the review. And on a study where the codeframe genuinely must not move, a strict/locked mode that blocks new codes from being added at all, so a coder under deadline pressure can't quietly widen a tracking study's categories without anyone deciding to.
Frequently asked questions
Is there a real difference between a codeframe and a codebook?
In careful usage, the codeframe is the category list; the codebook is that list plus the definitions, examples and edge-case rules that make it applicable by more than one person. In everyday use among researchers, the two words are interchangeable — both mean "the thing I code responses against."
How many codes should a codeframe have?
There's no fixed number — it follows from how varied the responses actually are, not from a target. What's worth watching is the shape: a handful of codes each covering 15%+ of responses, a longer tail of smaller ones, and an "Other" bucket in the single digits. If "Other" is over 15-20%, the codeframe is missing something real.
Can the same codeframe be reused across survey waves?
Yes, and on a tracking study it should be — that's what makes wave-over-wave comparison mean anything. The practical version is keeping one shared codebook file and reusing it for every wave rather than rebuilding it, so a code doesn't quietly get renamed or split between Q1 and Q2.
What happens to a response that doesn't fit any existing code?
It gets flagged rather than forced into the nearest wrong category. From there it's either a genuine one-off (goes to "Other"), evidence the codeframe is missing a category (a new code gets added and the decision applied retroactively), or — on a locked tracking codebook — logged for the next redesign rather than acted on immediately.
Where this fits
A codeframe is the input; coding is the process of applying it to every response, and survey coding covers that process end to end. Commercial coding has its own constraints — deadlines, multiple coders, client deliverables — separate from the academic version of the same discipline, covered in open-end coding in market research. And once responses are coded, what a client actually reads is built on top of the codeframe: see cross-tabulation analysis for how that table gets built and read.
If you want to see a codeframe get generated rather than read about one, our free open-end coding tool takes up to 100 responses, proposes a codeframe from the actual text, and returns it as an Excel file — no account required.