Insights
Heuristic evaluation, usability testing and a UX audit answer different questions. Match yours to the method, the time it takes and what you get. Pick the method by the question you need answered.
Pick the method by the question you need answered. A heuristic evaluation checks an interface against established usability principles, quickly and without users. Usability testing shows whether representative users can complete specific tasks. A UX audit, when it is scoped well, explains why friction keeps returning and what to fix first. You may need two, in sequence.
Last reviewed September 29, 2026.
Should I get a UX audit, a heuristic evaluation or usability testing?
Start from the question on the table, not the name of the service. The last row matters as much as the others: some questions none of the three can answer.
| The question you have | Method that answers it | Time it takes | What you get |
|---|---|---|---|
| Does this screen or flow break well-known usability principles? | Heuristic evaluation | Usually the quickest. No recruiting; time grows with scope and the number of evaluators (three to five is the usual advice). | A list of problems, each tied to a principle and rated for severity. |
| Can the people who use our product complete this task, and where do they fail? | Usability testing | Usually the longest lead. Recruiting representative users, writing tasks, sessions of 60 to 90 minutes each, then analysis. | Observed failures and the circumstances around them. Evidence of what happened, not of how often. |
| Will this design work before we build it? | Usability testing on a prototype, often after a quick heuristic pass | As above, plus the time to make a prototype detailed enough to test the risky part. | A tested direction, the changes it needs and what remains unproven. |
| Why does the same friction keep coming back, and what should we fix first? | A UX audit whose scope includes cause, not only symptoms | Depends on how much evidence has to be gathered: product access, analytics, support records, people who know the workflow. | Findings tied to evidence, likely causes, and an ordered set of actions with owners. |
| How many of our users hit this problem? | None of the three. Use analytics or a quantitative study. | Depends on data already collected; a quantitative test needs far more participants than a qualitative one. | A frequency or rate, with a margin of error. |
What each one is
Heuristic evaluation. Evaluators judge a design against a set of usability guidelines, commonly Jakob Nielsen's ten usability heuristics, the set NN/g recommends, such as visibility of system status and error prevention. No users take part. Nielsen Norman Group (NN/g) recommends three to five evaluators working independently, because each one misses problems: averaged over six of Nielsen's early projects, a single evaluator found only about 35% of the problems. The method works on a prototype or even paper, before anything is built.
Usability testing. A facilitator asks participants to perform realistic tasks with a product or prototype and watches what they do (NN/g's definition). The evidence is behavior, not expert opinion. A small qualitative round is built to find problems. Nielsen's often-quoted model says five participants find about 85% of the problems, but it assumes each user runs into about 31% of them, and it applies per user group: with two distinct user groups, Nielsen suggests three to four users from each, and three from each when there are three or more groups. Typical sessions run 60 to 90 minutes. A round like that cannot tell you how many of your users will hit a problem; NN/g notes that estimates are usually precise enough only with 40 or more participants.
UX audit. There is no standard definition. NN/g sells its expert design review as a UX audit service: a specialist inspects the interface and rates findings by severity. NN/g describes an expert review as a more general version of a heuristic evaluation, drawing on wider guidelines and the reviewer's experience. Other agencies use the name for something much broader; one agency's description adds stakeholder goals, heatmaps and session replays, workflow mapping, accessibility and competitor benchmarking. Two quotes for a "UX audit" can describe different work, so ask what evidence goes in and what decision comes out.
What the research says about inspecting versus watching
A common belief is that testing with users always beats expert inspection, which only finds cosmetic issues and false alarms. The evidence is more mixed than that.
- The older view. Research from the early 1990s, summarized by the Usability Body of Knowledge, found heuristic evaluation may surface more minor and fewer major problems than a think-aloud test (Jeffries and Desurvire, 1992).
- A direct comparison. In CUE-4, 17 professional teams evaluated the same hotel website at the same time in 2003; nine ran usability tests and eight used their preferred inspection method. Its authors, Rolf Molich and Joseph Dumas, report that inspections by experienced practitioners were comparable to usability tests in the pattern of problems they found, and the expert reviews took significantly less time. Many teams produced results that could drive a round of design changes in under 25 person-hours.
- Evaluators disagree, whatever the method. A review of eleven studies by Morten Hertzum and Niels Ebbe Jacobsen found that the average agreement between any two evaluators using the same method on the same system ranged from 5% to 65%, and that this held for think-aloud testing as well as heuristic evaluation. In CUE-4, even the last team found new, valid problems.
- Domain knowledge helps. Evaluators with both usability and domain expertise found the most problems in Nielsen's 1992 work, as the same summary records.
Two cautions on reading this. CUE-4 looked at one consumer website with experienced teams, and only one of its eight inspection teams ran a formal heuristic evaluation; it is neighboring evidence for a complex B2B product with specialist roles, not direct evidence. And "comparable pattern of problems" does not mean the two methods find the same problems.
Our reading: the choice between inspecting and watching matters less than buyers tend to assume. Who does the work, what they know about your domain and which tasks they cover matter at least as much. Any single evaluation is a sample of your product's problems, not an inventory of them.
What each one cannot tell you
A heuristic evaluation cannot tell you whether your users actually struggle with a violation, or what they struggle with that no principle predicts. NN/g is plain that heuristic evaluations cannot replace user research, and notes that testing can surface issues an expert would not think of because the real audience has specific knowledge or needs. In B2B software, much of the difficulty sits in business rules, roles and exceptions, which a reviewer without domain context may not see.
Usability testing only covers the tasks you wrote and the people you recruited. Change the tasks or the facilitator and new problems appear. A small study can show a serious failure without telling you how widespread it is. And a participant completing a task does not prove they understood it or could recover from a mistake without help. For specialist B2B roles, finding representative participants can be the slowest part; a 2003 NN/g survey of recruiting agencies found high-end professionals cost about twice as much to recruit as average consumers.
A UX audit is only as good as the evidence and the sample behind it. If it is a heuristic evaluation under another name, it inherits the limits above. If it ends in a long list of findings with no causes, owners or order, it is unlikely to settle the disagreement that prompted it.
None of the three proves a product is ready to release. That needs its own checks of security, performance, permissions and operations.
A sensible order when you are not sure
These are judgments, not rules, but they follow from what each method can and cannot establish.
- Early design, nothing built. Run a quick heuristic pass on the prototype to clear obvious problems, then test the riskiest flow with representative users. Testing a prototype full of known violations spends participants' time on problems an expert could have caught.
- A live product where the team agrees what is wrong with one flow. Test that flow. You need to see where people fail, not another opinion.
- A live product where friction keeps returning and the team disagrees about why. Start with an audit that looks for cause. Repeated friction can point to accumulated UX debt or to an unresolved product rule rather than a single screen. A usability test may be one of the audit's recommendations, aimed at the question the audit could not settle.
- You need to know how often something happens. Use analytics or a quantitative study; none of the three methods here is built for that.
- You already know the change you need. Skip the diagnosis and commission the work.
Questions to ask before you commission any of them
- [ ] How many evaluators will review the product, and will they work independently before comparing notes?
- [ ] What do they know about our domain, users and business rules, and how will they learn what they don't?
- [ ] For testing: who are the participants, how are they recruited, and do they match each role that uses the product?
- [ ] Which tasks or workflows are covered, and why those?
- [ ] Will the report separate what was observed from the explanation for it?
- [ ] How are findings ranked, and does the ranking reflect consequence and strength of evidence rather than an arbitrary score?
- [ ] Does each finding have a recommended action and someone to own it?
- [ ] What does the report say it cannot conclude, given the sample?
Where Tcules fits
Tcules' UX and Product Audit is built for the fourth row of the table: friction that keeps returning and a team that disagrees about the cause. It follows a representative workflow across the roles, states, content and exceptions involved, and compares product behavior with the intended work and business rules. It separates what was observed from plausible causes and states where the sample cannot support a wider conclusion. The result is an ordered set of decisions, each with an owner, a recommended action and a validation step. It does not produce a heuristic score; Tcules avoids arbitrary scoring systems and percentages that imply more certainty than the evidence supports. Your team can act on the findings without Tcules, and this guide covers how to use them.
If your question is whether people can use a design you have not built yet, that is Prototyping and Usability Testing: a prototype detailed enough to expose the uncertainty, tested with representative users.

