Lists of research methods are easy to find and they solve a problem almost nobody has. Teams are not stuck because they have never heard of a diary study. They are stuck because the question they were handed does not name its own shape, and the wrong method answers it confidently.
So this starts from the question rather than the menu.
Four shapes of question
Almost every research request is one of these, and identifying which one you have eliminates most of the method list immediately.
How many, how much
Prevalence, size, distribution, comparison between groups. Answerable only by asking enough people in a structured way, and the whole design question is sample and sampling.
Why, how
Mechanism. What led to a decision, how a process actually runs, what somebody was thinking. Answerable only by talking to people or watching them, and the whole design question is who you talk to.
Can they
Whether somebody can complete a task, understand a message, or find a thing. Answerable by observation under realistic conditions, and small numbers work because failures repeat.
Would they choose it
Preference and behaviour under a real trade-off. The hardest shape, because asking directly produces an answer with a weak relationship to what people actually do.
Asking a survey why something happened produces an answer. It just will not be the reason.
The methods, and what each cannot do
| Method | Answers | Cannot answer | Typical n |
|---|---|---|---|
| Survey | How many, how much | Why, and anything about future behaviour | 200 – 2,000+ |
| Depth interview | Why, how, what happened | How common any of it is | 8 – 20 |
| Diary study | What happens over time, in context | Anything needing scale | 15 – 40 |
| Usability test | Can they complete it | Whether they would want to | 5 – 8 per audience |
| Concept or creative test | Was it understood, is it believed | Whether it will sell | 50 – 150 per variant |
| Field observation | What people actually do | What they were thinking | 6 – 15 |
| Controlled experiment | Would they choose it | Why they chose it | Depends on effect size |
The right-hand column is the useful one. Every method on that list is routinely asked to answer something in its own "cannot" column, and it always produces output — which is the problem, because output looks like a finding.
The mistake that accounts for most wasted research
Using a survey to answer a "why" question.
"Why did you cancel?" in a multiple-choice list returns the option that is easiest to select and hardest to argue with. "Too expensive" is the standard answer and it is usually a summary of something more specific: a feature that stopped working, a competitor's offer, a change of role, an invoice that arrived at a bad moment. None of those are available from the list, and all of them are actionable in a way the list answer is not.
The reason teams do it anyway is availability. The survey tool is already licensed, the list of customers already exists, and twenty interviews look expensive next to a free email send. The interviews are cheaper than the quarter spent acting on a wrong reason.
What people say against what they do
The distinction runs through every method and it is not a reason to distrust customers.
People are reliable about what happened. Ask about a specific past event — what they did, in what order, what went wrong — and the account is usually good, particularly if you anchor it to a recent instance rather than to general habit.
They are much less reliable about what they would do. Stated intent is shaped by how the question was asked, by what sounds sensible to say, and by a version of the future that has none of the friction the real one will have. This is why purchase-intent scores are the most quoted and least predictive number in the field.
The practical rule: for behaviour, design a situation rather than a question. An experiment with a real choice, a landing page with a real button, a task with a real outcome. Where that is impossible, treat stated preference as a hypothesis rather than a measurement. What creative pre-testing can and cannot establish works through the version of this that comes up before a campaign launches.
Sample size follows the question
There is no general answer, and the general answers people quote come from a different question than theirs.
Discovering problems
Five to eight people per distinct audience. Usability failures repeat quickly, and the eighth participant rarely shows you something the first five did not. Adding people here buys reassurance, not information.
Understanding mechanism
Eight to twenty interviews, stopping when new conversations stop producing new categories. If the twelfth interview is still surprising you, the audience is broader than the study assumed.
Directional comparison
A few hundred responses, sampled from real traffic or a matched panel rather than from a hand-built list.
Detecting a small difference
Thousands. The requirement scales with the square of the effect, so halving the difference you want to see quadruples the sample. The arithmetic behind that is worth reading before promising a two-point measurement.
The two failure modes are symmetrical and both common. Generalising from five people to a percentage is one. Demanding statistical significance from a discovery study is the other, and it produces expensive research that answers a question nobody asked.
Recruitment is what actually costs
Method choice affects the budget far less than who has to be in the room.
Incidence dominates everything. If one person in fifty qualifies, the study pays for forty-nine screeners to reach one respondent, and that arithmetic swamps the difference between a cheap method and an expensive one. A general consumer study and a study of people who manage a specific type of contract can use identical designs and differ tenfold in cost.
Market and language narrow the pool further, and the curve gets steep quickly. Professional qualification narrows it further still. How each layer of targeting changes both the pool and the rate sets out where each requirement starts to hurt, and why incidence dominates recruitment cost works through the numbers.
The incentive is the other lever, and it is usually set by precedent rather than reasoning. Too low and only the extremes respond; too high for the task and qualifying becomes the goal. How to set it and what each form of payment does to who accepts covers the arithmetic.
Sequencing beats method selection
Most well-run programmes alternate, and the order matters more than the choice.
Qualitative first, to learn what to ask. Eight conversations tell you which questions are worth putting to two thousand people and, more usefully, which of your assumed answer options nobody would ever pick.
Quantitative second, to find out how widely it holds. This is where an interesting anecdote becomes a prevalence you can plan around, or fails to.
Then qualitative again on the surprising cells. A quantitative result you did not expect is a question, not a conclusion, and the cheapest way to understand it is to talk to a few people in that group.
Teams that run only the second step have numbers nobody can explain. Teams that run only the first have vivid stories with unknown reach. Both are common and both are recoverable by adding one step rather than by changing method.
The failure modes worth naming
Five, and they are recognisable once you have seen them.
Leading questions. "How useful did you find the new dashboard?" presumes usefulness and gets it. The neutral version asks what they used it for and what happened, and it produces a different answer often enough to be worth the discipline.
Convenience samples described as representative. The customers who agreed to a call are not a random draw from your customers — they are the engaged ones, and every finding carries that bias whether or not the write-up mentions it.
Asking customers to design. People are excellent witnesses to their own problems and unreliable architects of solutions. "What would you like us to build" produces a faster horse; "walk me through the last time this went wrong" produces the brief.
Analysis by whoever ran the study. The person who conducted twelve interviews remembers the vivid ones, and recency and rapport shape a summary more than anybody expects. Have somebody else read the raw material before the conclusions are fixed.
And treating an absence of evidence as evidence. A study that found nothing frequently means the sample was too small to find it, not that the effect is absent. Those are different results and only one of them justifies a decision.
When there is no budget
Three things cost almost nothing and are consistently underused.
Read what you already have. Support tickets, cancellation reasons, review text, sales call notes, search queries on your own site. Most companies hold more unread qualitative data than they could collect in a quarter.
Talk to five customers. Not a programme, not a screener, just five conversations about a specific recent event. It is the highest-return research activity available and the one most often postponed for a proper study that never gets commissioned.
And watch somebody use the thing. Six people attempting a real task will find the majority of what is wrong with it, and what structured usability testing adds on top of that is a matter of rigour rather than of discovering something entirely different.
The cheapest research available to most companies is reading the feedback they already collected and never opened.
Where multi-market changes the design
A method that works in one market does not transfer by translation.
Scale use differs, so satisfaction and agreement scores are not comparable across countries without adjustment that few programmes make. Register differs, so a verbatim rendered into English loses the difference between mild annoyance and real anger. Willingness to criticise differs, and in some markets a polite answer to a direct question is close to uninformative.
The workable approach is to compare within a market over time rather than between markets at a point in time, to have open-ended responses coded by somebody native to each market, and to treat each market as its own sample rather than as a slice of a total. Which populations a feedback programme never reaches covers the related problem of who is missing from the data you already hold.
Choosing, in one paragraph
Name the question and its shape before naming a method. If it is how many, plan for sample and sampling. If it is why, plan for who you talk to and stop worrying about numbers. If it is can they, six people will do and you should watch rather than ask. If it is would they, build a situation with a real choice in it, because the survey version of that question has been answered wrongly for decades. Then check feasibility and incidence before anything else, since the study you cannot recruit for is the only one guaranteed to fail. How a brief becomes reserved capacity and verified responses describes what that looks like when the people you need are not on anybody's panel.
Where the answer is that you need a survey, how to design and sample one covers question types, wording and the sampling decisions that carry most of the error.