Worked example
Two structures for a connected-car app, and a comparison that was only partly like for like
The question. Companion apps for cars usually group features by vehicle system: climate, security, charging. Should a new app follow that convention, or organise features by whether you are near the car or away from it?
Evidence already available. An audit of existing companion apps gave a list of 24 features. None of the audited apps organised features by whether you are near the car.
Card sort. In an open card sort, the first 60 US adult participants sorted the 24 features. Most grouped them by vehicle system. About a quarter grouped them by a different rule: near the car, or away from it. The sort produced two plausible grouping schemes; it could not tell which would be easier to navigate.
Two candidate structures. The team built one structure on each organizing principle: A followed the vehicle-system convention, B the near-or-away principle. Both held all 24 features.
Tree test. Each structure was tested with a separate group of participants on eight tasks with the same intent, such as starting the car remotely or finding where it is parked. After excluding internal sessions, one outlier, and one participant who completed both tests, 18 participants per structure remained, with 144 attempts each.
Result. Overall success was 60% for A (87 of 144 attempts) and 70% for B (101 of 144). Direct success was identical: 63 of 144 for each.
What the comparison could not separate. Three of the eight tasks were worded differently in the two tests. In A, one read "make sure it's protected while you're gone"; in B, "turn on the security monitoring". Another read "make sure you have enough range" in A and "check your fuel level" in B, where B's correct item was labelled Fuel Level. These three tasks include the two largest gaps between the structures, and B's lead overall comes from them: 44 of 54 successes for B against 26 of 54 for A.
On the five tasks worded identically, the structures were close: 61 of 90 successes for A, 57 of 90 for B, and 42 of 90 direct for each. The clearest matched difference was directness on one task, unlocking the car from home: 3 of 18 went straight to the answer in A, 11 of 18 in B, with similar overall success.
Where the overall gap came from
5 tasks worded the same for both
A · Vehicle system
61/90
B · Near or away
57/90
3 tasks worded differently
A · Vehicle system
26/54
B · Near or away
44/54
All tasks
A · Vehicle system
87/144
B · Near or away
101/144
Matched: A 61/90 · B 57/90. Mismatched: A 26/54 · B 44/54. All tasks: A 87/144 · B 101/144
| Tasks | A · Vehicle system | B · Near or away |
|---|---|---|
| 5 tasks worded the same for both | 61/90 | 57/90 |
| 3 tasks worded differently | 26/54 | 44/54 |
| All tasks | 87/144 | 101/144 |
What it supports. Matched-task results did not favor one structure overall. B had higher Direct success on one remote task. Both remained plausible candidates.
What it does not support. The results do not support the claim that B is better overall: the overall gap comes from tasks that were not worded the same. With 18 participants per structure, differences of this size on single tasks remain uncertain. Tree testing does not measure live-interface performance.
What would settle it. Re-test the three mismatched tasks with one wording for both structures, written around the need rather than either label, and with more participants per structure if the choice matters.