Open-end coding in commercial market research is a different discipline from qualitative coding as it is taught. Both turn text into categories. Almost nothing else is shared: not the timescale, not the unit of analysis, not what counts as a finished job.
The distinction matters because most of the written guidance comes from the academic tradition, and a research team that follows it will produce something their client cannot use.
What makes commercial coding different
The output is a column, not an argument. Academic coding produces an interpretation defended in prose. Commercial coding produces a variable that has to cross-tabulate against age, region and brand usage. If it cannot be counted and banner-tabbed, it is not finished.
The codeframe is negotiated with the client. It is a deliverable in its own right, often approved before coding begins, and frequently inherited from a previous wave that somebody else ran. You rarely get to invent it.
Multi-coding is the norm, not an exception. A single verbatim routinely carries three codes. Frameworks that assume one label per unit do not survive contact with "great range, awful staff, and parking is impossible".
The deadline is measured in days. Not because commercial research is careless, but because fieldwork closes on a Friday and the debrief is on Wednesday. Any method that cannot absorb four thousand verbatims inside that window is theoretical.
The codeframe is the instrument
Everything downstream inherits the frame's flaws, which is why it deserves more attention than the coding itself.
A frame that works has codes at one level of abstraction — mixing "price" with "the loyalty discount expired without warning" guarantees inconsistent application. It separates topic from sentiment, so you can still count mentions of delivery independently of whether people liked it. It has explicit rules for the boundary cases, because "does an unprompted competitor mention count as a criticism?" will come up on verbatim forty and again on verbatim nine hundred, and two coders will answer differently unless it is written down.
And it reserves a residual code that somebody actually reads. The unplaced responses are where the next wave's codes come from.
Waves change everything
In tracking work the frame stops being a working tool and becomes the measurement instrument. Wave-over-wave comparison is only valid if the same thing was measured the same way, which means the frame is frozen and changes to it are events with consequences.
The practical discipline: new themes get new codes, never redefinitions of existing ones. A code that changes meaning mid-track produces a trend line that is wrong without being visibly broken — the worst failure mode available, because it looks like a finding.
When a genuinely new theme appears — a product recall, a competitor launch — the honest move is a new code with a start date, and a footnote on every chart that spans the boundary.
What quality control looks like here
Academic practice reaches for inter-coder reliability statistics. They are useful, but on a commercial timeline the more informative checks are cheaper:
- Residual rate per question. A question with a high residual has the wrong frame, and the number tells you which one to fix.
- Code frequency distribution. A code applied to sixty percent of responses is doing no work. A code applied twice may be noise, or may be the finding — read them.
- Blind double-coding of a sample. Two hundred responses, two coders, compare. The point is locating which codes are ambiguous, not producing a coefficient.
- Reading the verbatims behind the top code. The most-applied code is the one most likely to have quietly become a catch-all.
The deliverable
A finished coding job hands over more than coded data. It hands over the codeframe as an artefact the client can inherit, the residual responses with their counts, a note of every boundary rule that had to be decided, and data in a format that goes straight into tabulation — which in practice means an SPSS file with value labels intact, not a spreadsheet somebody has to reshape.
That last point is where a great deal of time disappears in commercial research: coding that is methodologically sound and then arrives in a shape the tab house has to rebuild.
The mechanics of building and applying a frame at scale are set out in our survey coding guide, and the specific ways automated coding fails in AI open-end coding.