Usability Testing, Card Sorting, and Contextual Inquiry: How to Choose the Right UX Research Method
October 08, 2026
Introduction
Last spring, a product manager asked me which research method she should use. Her team had a new dashboard, a shrinking number of active users, and three different theories about the cause. She wanted a single answer. I told her the honest version: the method depends on the question, and the wrong method will give you a confident answer to the wrong question.
That conversation is more common than people admit. Teams often reach for usability testing because it feels concrete, when the real problem is that they don't know what users need. Others run interviews when they already know what the problem is and only need to check whether a fix works. In this post I'll walk through the main methods, explain what each one can and can't tell you, and show how to combine them without wasting budget. Along the way I'll point out where the choice matters for the kind of UX research services your team might buy.
Start with the question, not the method
Every research method answers a specific kind of question. Before picking one, write down what you need to know in a single sentence. "Why do people stop during onboarding?" calls for observation and task-based testing. "What should we call this section of the menu?" calls for card sorting. "What do small clinics need from scheduling software?" calls for interviews and possibly contextual inquiry.
A useful test is to ask what you would do differently depending on the answer. If the result wouldn't change a decision, the study is probably the wrong one, or it's answering a question you didn't need to ask. I've seen teams spend three weeks on usability testing when the answer they needed was already in their support tickets.
It also helps to separate generative questions from evaluative ones. Generative questions ask what to build. Evaluative questions ask whether what you've built works. Most of the methods below fall clearly into one group or the other, and the mismatch between a question and a method is one of the most common reasons research disappoints.
Usability testing: does this design work?
Usability testing puts a design in front of real users and asks them to complete specific tasks. The researcher watches what happens, notes where people hesitate or fail, and records the reasons they give. The output is usually a task success rate and a list of observed problems, ranked by how often they occurred and how badly they affected the task.
The method is strong on specifics. If participants can't find the export button, you'll see the moment it happens. If they click the wrong menu item three times, you'll see that too. Usability testing is also relatively quick. A round can run in 1 to 3 weeks, depending on recruitment.
Its limit is scope. Usability testing tells you whether a particular design supports a particular task. It doesn't tell you whether the task was the right one to support. A flawless checkout flow can still sit on top of a product nobody needs. That's why I push teams to run usability testing after they've already done some generative work, not as a substitute for it.
On the question of sample size, the Nielsen Norman Group's article on testing with five users remains the clearest explanation I know. The core idea is that a single round with five participants surfaces most of the serious problems, and several small rounds beat one large round. That logic is why we usually run usability testing in iterations rather than as one big study.
Card sorting: does the structure match how people think?
Card sorting asks participants to group a set of topics into categories that make sense to them. You can run an open sort, where people create and name their own groups, or a closed sort, where they place items into predefined groups. Open sorts reveal the language people use. Closed sorts test whether your existing categories hold up.
A card sort is especially valuable before a redesign of navigation or information architecture. Teams often argue about labels, such as whether a feature belongs under "Billing," "Account," or "Settings." A card sort gives the argument a reference point. The Nielsen Norman Group's overview of card sorting explains the method and its variations in detail.
The usual follow-up is a tree test, which asks participants to find items in a proposed structure without any visual design getting in the way. The combination tells you whether the structure works in principle and whether the labels are clear. In our programmes, a card sort and tree test typically use 30 to 50 participants, since the patterns need volume to stabilize.
Card sorting has a blind spot. It tells you how people group concepts, not how they behave under pressure. A participant might sort "Refunds" under "Payments" in a calm session, then search for it under "Support" when a customer is angry. That's why card sorting works best alongside observation.
Contextual inquiry: what do people actually do?
Contextual inquiry takes research into the place where the work happens. The researcher observes a person doing their real task and asks questions as the work unfolds. The method is built on the idea that people are poor witnesses to their own routines. They skip steps, they forget workarounds, and they describe an idealized process when asked to explain how they work.
I find contextual inquiry most useful for complex B2B products and for work that happens under constraints. A nurse checking medication schedules, a warehouse supervisor reconciling stock, or a finance analyst closing the month will all describe their work differently in an interview than they perform it. Observation shows the real sequence, the interruptions, the paper printouts, and the second monitor with a spreadsheet nobody mentioned.
The trade-off is cost and time. Contextual inquiry means travel, scheduling around real work, and careful notes. We usually observe four to six users, which is enough to see the patterns without drowning in data. For a project with a tight budget, a good approach is to run a few contextual inquiry sessions with the most important user group, then rely on interviews for the rest.
It's worth knowing the difference between contextual inquiry and the broader term ethnographic research. Ethnographic research studies a culture over a longer period, often weeks or months. Contextual inquiry is a shorter, more focused version applied to a specific task. Many product teams borrow the technique without the full ethnographic scope, which is perfectly reasonable if you're clear about what you can and can't conclude.
User interviews: why do people do what they do?
Interviews are the most common generative method, and they're often underused because they seem simple. A good interview is a structured conversation that moves from the participant's context to their goals, their workarounds, and their emotional responses. The skill is in asking about specific recent events rather than hypothetical future behaviour. "Tell me about the last time you exported a report" produces better data than "Would you use an export feature?"
We usually run 10 to 15 interviews per user group. Patterns start repeating after that, and new interviews add diminishing returns. The most useful finding in an interview is often an unexpected detail, such as a participant who prints every document before reviewing it. Those details rarely show up in a survey, and they can reshape the whole product direction.
Interviews are not the same as focus groups. In a focus group, participants influence each other and the loudest voice can dominate. One-to-one interviews are better for understanding individual behaviour, which is what generative research needs.
When to combine methods
The strongest research programmes combine methods, and the order matters. A common sequence runs interviews first, then card sorting to test the structure those interviews suggested, then usability testing on the resulting design, and finally a round of evaluation after launch. Each step narrows the question and uses the previous step's findings.
Here's an example from a project type I see often. A team wants to redesign the dashboard for a B2B tool. Interviews reveal that users check three specific numbers every morning and rarely look at anything else. A card sort then shows that users expect those numbers under the label "Today," not "Overview." Usability testing on the new layout shows that one of the three numbers still gets missed because it sits below the fold on laptop screens. Three methods, three distinct findings, and each one changed a decision.
Compare that with a team that runs only usability testing on the existing dashboard. They'll learn which buttons confuse people, but they won't learn that the dashboard is showing the wrong information. That's the core reason I recommend a full programme when the stakes are high. The UX audit services we offer can catch many surface issues quickly, but research answers the deeper question of whether the product is solving the right problem.
Choosing between a UX audit, usability testing, and full research
Teams often confuse three options: a UX audit, a usability testing round, and a complete research programme. They differ in cost, timing, and the kind of answer they produce.
Option | Main question | Typical timeline | What you get |
UX audit | Where does the current product fail? | 2 to 3 weeks | Prioritized fix list with severity and effort |
Usability testing | Can people complete these tasks? | 1 to 3 weeks per round | Task success rates and observed failures |
Full UX research programme | What should we build, and does it work? | 4 to 8 weeks | Interviews, card sorts, personas, journey maps |
If your product already exists and you suspect it has obvious problems, an audit is usually the cheapest way to start. If you have a design and want to know whether specific tasks work, usability testing is the right fit. If you're not sure what to build, or the product strategy is in question, the full programme is the better investment. The usability testing services page explains how the rounds are structured if you already know you need that step.
Common mistakes to avoid
The first mistake is testing too late. By the time a design is polished and handed to engineers, changing it feels expensive, so teams test to confirm rather than to learn. Test early, with rough sketches if necessary.
The second is recruiting the wrong people. A study that includes power users and first-time visitors in the same sample will produce conflicting findings. Decide which user group matters for the decision, then recruit for it.
The third is treating a single finding as proof. One participant struggling with a button is a signal, not a verdict. Look for patterns across participants before changing the design.
The fourth is skipping the synthesis. Raw notes don't change anything. The team needs a clear summary of what was found, how often, and what it means for the next decision.
The fifth is ignoring accessibility. Usability testing should include participants who use assistive technology when the product serves them. A product that works for a sighted mouse user can still fail a screen reader user, and the WCAG 2.2 guidelines from the W3C are the baseline I'd use for checking that.
What I'd do in your position
If I were starting with a product that has real traffic but unclear retention, I would begin with a short audit to find the obvious failures. Then I'd run interviews with the users who stopped, not only the ones who stayed, because the people who left often explain the problem best. Only after those findings would I decide whether a card sort or usability testing makes sense. That order keeps the budget focused on questions that change decisions.
For a new product, I'd flip the sequence. Start with interviews and contextual inquiry to learn what problem is worth solving. Use card sorts to shape the structure, then test concepts before any engineering starts.
Whichever path you choose, decide the question first and the method second. The method is the easy part once the question is clear.
If you're weighing these options for a specific product, our UX research services cover the full sequence, from interviews through evaluation, and we can scope a smaller engagement if the full programme isn't the right fit yet.
Frequently asked questions
What is the main difference between usability testing and card sorting?
Usability testing checks whether people can complete tasks with a specific design. Card sorting checks how people group information and what labels they expect. Usability testing answers "does this work?" and card sorting answers "does this structure make sense?" Teams often need both when they redesign navigation.
When should I use contextual inquiry instead of interviews?
Use contextual inquiry when the work is complex, happens in a specific physical or operational setting, or involves steps people might forget to describe. Interviews work well for understanding goals and attitudes. Observation works better for understanding real sequences and workarounds.
How many participants do I need for a card sort?
For a reliable open card sort, 30 to 50 participants is a common target. Smaller sorts can reveal broad patterns, but the results become less stable with fewer people. Closed sorts can work with fewer participants because the categories are predefined.
Is usability testing enough to decide what to build?
Usually not. Usability testing shows whether a design supports specific tasks, but it doesn't tell you whether those tasks are the right ones. Generative methods such as interviews should come first when you're deciding what to build.
How many users should I test in a usability round?
Five participants per round uncovers most of the serious problems in a design, according to the well-known Nielsen Norman Group finding. Running several small rounds, with changes between them, usually produces better results than one large study.
What does a contextual inquiry session look like?
The researcher watches a participant do their real work and asks questions as the task unfolds. Sessions typically last one to two hours, depending on the task. The researcher takes notes and sometimes photographs the workspace, with consent. The focus is on what people do, not on what they say they do.
Can I run these methods remotely?
Interviews, card sorts, tree tests, and most usability testing work well remotely. Contextual inquiry is often better in person, although remote versions exist for some digital workflows, where participants share their screen while working. The right choice depends on where the work happens.
How do I know which method to start with?
Write down the decision you need to make. If you're deciding what to build, start with interviews or contextual inquiry. If you're deciding how to structure navigation, start with card sorting. If you're deciding whether a design works, start with usability testing. If you're not sure what the problem is, a UX audit is often the fastest starting point.
How does a UX audit differ from usability testing?
An audit is an expert review of the product against usability principles and accessibility standards, and it doesn't involve recruited users. Usability testing involves real participants completing tasks. An audit is faster and cheaper, while usability testing provides evidence from actual users. Many teams use both, starting with the audit.
What is the biggest mistake teams make with UX research methods?
Testing too late is the most common mistake. When research happens after the design is finished and engineering has started, findings tend to confirm decisions rather than inform them. Starting research early, even with rough sketches, gives the team room to act on what they learn.