Tree testing example

Tree testing an e-commerce admin: relabel or restructure?

By UXbeam, information architecture tools and services since 2021 · Updated August 9, 2026

One rename raised Direct success 19 points. The broader revision moved overall Direct 3 points.

Feature-dense products develop the same navigation problem: features with several plausible homes. Should automated review requests live with customer communications, with orders, or with analytics? UXbeam tree tested three versions of a seller-side e-commerce admin, Initial, Iteration 1, and Iteration 2, against the same nine tasks.

Automatic review request
Initial 59 panel sessions
Iteration 1 67 panel sessions
Iteration 2 (directional) 38 panel sessions
What was tested

Three versions of one seller admin: trees of 60 to 80 items, nine tasks.

With whom

59, 67, and 38 panel sessions across the three versions.

What moved

Renaming one menu lifted its task from 27% to 46% Direct. The full revision moved overall Direct only from 30% to 33%. Iteration 2 (directional) reached 48%.

Where tree testing fits

1Feature inventoryfour seller apps
2Initial IAthe first tree
3Testnine tasks per version
4Iteratechange what the clicks show
5Test againIteration 2
6Move into designwireframe with a validated tree

The setup: a seller admin based on Shopify, Etsy Seller, WooCommerce and Ecwid Android apps

The tested product is a seller-side admin: the screens a seller uses to run a store, not the storefront a buyer sees. It covers orders and shipping, listings, inventory, customer communications, promotions, and sales analytics, in trees of 60 to 80 items. The feature set drew from four seller apps, each serving sellers counted in the hundreds of thousands or millions.

One of those jobs is collecting product reviews. After a buyer receives an order, the store can automatically email them asking for a reviewWooCommerceWooCommerce DocumentationRemind customers to leave a reviewAutomateWoo docs · woocommerce.com. The reviews it gathers shape how future shoppers judge a product: the first few displayed reviews measurably raise purchase likelihoodSpiegelSpiegel Research Center, NorthwesternHow Online Reviews Influence Sales2017 · about 270% higher conversion at five reviews vs none. A seller who cannot find where to turn that on never reaches any of it.

Where does this live?

Customer communications ? Orders ? Sales analytics ?
Three plausible homes for one feature.

Admin findability is a daily cost. Sellers arrive with a job in mind: ship the order, check stock, answer a buyer. The work repeats, so every menu that makes someone stop and decode adds up; Nielsen Norman Group's intranet research calls navigation labels "the most costly words in the company"NN/gNielsen Norman GroupIntranet Information Architecture (IA) MethodsJakob Nielsen · 2007 · nngroup.com. A tree test measures that cost before the structure ships.

Nine tasks, all from a seller's normal week:

  1. Pause notifications for the weekend.
  2. Share a new listing on Facebook.
  3. Find the closest post office to ship a new order.
  4. Set up an automated email asking buyers to review their order.
  5. Find which items are running low in inventory.
  6. Send a coupon to the most loyal buyers.
  7. Rebuild a dated store website.
  8. Open a second store to sell tutorials.
  9. Add a product video to a listing.

Participants saw only the labels and hierarchy, one task at a time. Each attempt ends Direct (right place, no backtracking), Indirect (right place after a detour), or Elsewhere (a confident finish somewhere wrong). The method is covered in tree testing, end to end; for phrasing tasks like these, see the Job-Shaped Task.

Iteration 1 renamed, split, and rebuilt three areas

Customer management center became Customer Notifications. The combined Manage orders/sales section split in two. The review-email area was rebuilt inside the renamed communications section. Seven positions left the tree, twelve arrived.

One setting, three menu names
Initial
Customer management center
  Notification settings
    Pause
Iteration 1
Customer Notifications
  Notification settings
    Pause
Iteration 2
Notifications
  Notification settings
    Pause

The path to the same setting in each version. Only the menu name changes.

1Finding

Iteration 1 moved overall Direct only 3 points

Initial and Iteration 1 ran head-to-head on the same nine tasks: 59 and 67 panel sessions.

Overall Direct: 30% on Initial (144 of 477), 33% on Iteration 1 (187 of 568).

Three points. The average hides one large win, several flat tasks, one scoring artifact (method note), and two problems the revision left untouched. Tree testing separates changing a lot from changing the right thing.

Overall Direct by version
Initial
30% · 144 of 477 · 59 sessions
Iteration 1
33% · 187 of 568 · 67 sessions
Iteration 2
48% · 158 of 328 · 38 sessions

Share of all attempts ending Direct. Iteration 2 ran later with some tasks reworded; read it as directional.

2Finding

One rename raised Direct success from 27% to 46%

The task read the same in both versions: pause notifications for the weekend. Only the tree changed.

On Initial, the setting lived under Customer management center; first clicks split three ways, and 27% ended Direct (16 of 59). On Iteration 1 it sat under Customer Notifications; 38 of 67 clicked it first, and Direct reached 46% (31 of 67).

The click pattern gives a strong clue: people scan for the word they already have in mind. Customer management center offered nothing to match. Customer Notifications did.

First clicks on the notifications task

Initial  first clicks split three ways:

Add new
22 of 59
Customer management center
20 of 59
Manage orders/sales
10 of 59

Iteration 1  38 of 67 first clicks on the renamed menu:

Customer Notifications
38 of 67
Add new
21 of 67

Iteration 2  33 of 38 first clicks on Notifications:

Notifications
33 of 38

The rename concentrated first clicks. Task wording was identical in Initial and Iteration 1.

3Finding

Two tasks stayed below 15%

Two tasks barely moved, and neither looks like a label problem.

The coupon task: 2% Direct on Initial (1 of 52), 6% on Iteration 1 (4 of 63). Both versions filed it under sales analytics; most first clicks went to communications, because sending a coupon reads as a message to buyers.

The inventory task crept from 12% (6 of 52) to 14% (9 of 63); Iteration 1 buried the alert one level deeper than most people searched.

The review-email task did worst: 9% Direct on Initial (5 of 54), 87% Elsewhere (47 of 54). No version cleared about 30% on it, and Iteration 1's gain partly reflects a scoring change (method note).

4Finding

48% chose the same alternative destination

Elsewhere is the most useful outcome in a tree test, because failures concentrate.

The initial tree held two nearly identical labels in the same menu: Auto send review request and Auto Sent review requests. Only the second counted as correct. 26 of 54 sessions (48%) finished at the first.

They were not lost; they picked the label the tree should probably keep. A concentration like that names the fix: merge the twins, or accept the one people choose. The share-listing task showed the same pull in miniature.

Where the review-email task ended (Initial)
Ended Elsewhere
47 of 54 (87%)
of which at "Auto send review request"
26 of 54 (48%)

Only "Auto Sent review requests" was marked correct. Its near-twin, one line away in the same menu, collected almost half the sample.

5Retest

Iteration 2 reached 48% Direct

Iteration 2 restructured around the behavior: top-level menus took the plain names people kept matching (Notifications, Orders, My stores, Sales analytics), inventory alerts moved under Sales analytics, and the review label people kept choosing, Auto send review request, was marked correct.

Across 38 panel sessions, overall Direct reached 48% (158 of 328). Read it as directional: a smaller test, with several tasks reworded, so its numbers are not a controlled comparison.

Notifications reached 82% Direct (31 of 38), with 33 of 38 first clicks on the menu. Inventory reached 41%. Coupon jumped to 54%, but its wording changed too, so part of that jump belongs to the task. Review email stayed unresolved.

Direct success on six tasks
InitialIteration 1Iteration 2 (directional)
Notifications
27% → 46%82%
Share a listing
21% → 38%38%
Post office
44% → 39%30%*
Review email
9% → 30%22%
Inventory
12% → 14%41%
Coupon
2% → 6%54%

Percent Direct per task. Filled dots: Initial and Iteration 1. Dashed hot-pink: Iteration 2, directional. † scoring or task wording changed (method note). ‡ label echoes the task wording; see "When to relabel, when to restructure". * task reworded in Iteration 2.

When to relabel, when to restructure

Four things, each visible in the data above.

Relabel when the label is the problem. The notifications feature sat in an acceptable place behind a name nobody matched. One rename moved first clicks from a three-way split to a majority, and the task from 27% to 46%.

Restructure when people keep choosing a different home. The coupon and review tasks stayed weak through Initial and Iteration 1, with the same destination pattern each time: people looking under communications, and finishing at the unmarked review label. The cleanest structural win is inventory: stuck at 12% and 14% until Iteration 2 moved it under Sales analytics and it reached 41%. For coupon and review, Iteration 2 is directional evidence, and review has still not cleared 30%.

Read first clicks and destinations before trusting a rate. The three-point overall move was the net of a 19-point win, several flat tasks, and a scoring artifact. First clicks showed why the win happened; destinations showed what to fix next. A top-line rate alone carries neither.

Treat your best number with suspicion. Iteration 2's winning label, Notifications, appears in the task instruction itself. UXbeam's leading-term check flags exactly this overlap, because people can match a word without navigating by meaning. We report the 82%; on our own results screen it would carry a flag. Spotting and rewriting that overlap is covered in leading tasks in tree testing.

Relabel or restructure?

Right area, wrong final labelRelabel
Same wrong home, repeatedlyRestructure
First clicks scatter everywhereCheck the task wording
A new version looks betterTest again before build

What UXbeam captures for every iteration

UXbeam is a tree testing and information architecture tool for testing navigation with participants and analyzing first clicks, paths and destinations. Each version in this study kept its own result set: task outcomes (Direct, Indirect, Elsewhere), first clicks, final destinations, complete navigation paths, and participant-level results. An iteration carries the tree and tasks forward, so the next test is set up in minutes rather than rebuilt. The charts on this page are that output.

Method note

  • Participants were US-based adults recruited through a paid online research panel.
  • Internal dry-runs were excluded from the reported counts.
  • Counts are panel sessions per version: 59 (Initial), 67 (Iteration 1), 38 (Iteration 2). A small number of panel IDs appeared more than once within or across versions, so no participant total is reported.
  • Per-task n varies because some participants stopped before finishing; each figure states its own n.
  • Duplicate submissions of the same task within a session were removed before reporting; the earliest submission was kept.
  • Initial and Iteration 1 were tested concurrently in April 2025; Iteration 2 in September 2025.
  • Task wording was identical between Initial and Iteration 1 except for small edits to the review-email, coupon, and tutorials tasks. Several instructions were reworded for Iteration 2, so its task-level numbers are directional.
  • Iteration 1 changed which destinations counted as correct for the review-email task (part of its gain is definitional) and removed one correct destination for the product-video task, so that task's apparent drop is partly a scoring change; it is excluded from improvement claims.
  • Rates are observed percentages of completed attempts. No statistical-significance claims are made.
  • This is UXbeam's own research on a client product, run to the same standard we recommend to teams.
See all nine tasks
Task Initial Direct Iter 1 Direct Iter 2 Direct Note
1 · Pause notifications27% (16/59)46% (31/67)82% (31/38)Identical wording in Initial and Iteration 1. Iteration 2's label echoes the task word; see the caution above.
2 · Share a listing21% (12/56)38% (25/65)38% (14/37)Steady pull toward the listing screen in every structure.
3 · Post office for shipping44% (24/55)39% (25/64)30% (11/37)Task reworded in Iteration 2.
4 · Automated review email9% (5/54)30% (19/63)22% (8/37)Iteration 1 changed which destinations counted as correct; part of the gain is definitional. Never cleared about 30%.
5 · Low inventory12% (6/52)14% (9/63)41% (15/37)Moved under Sales analytics in Iteration 2.
6 · Coupon for loyal buyers2% (1/52)6% (4/63)54% (20/37)Task reworded in Iteration 2; part of the jump belongs to the task.
7 · Rebuild store website51% (26/51)39% (24/61)54% (19/35)Task lightly reworded in Iteration 2.
8 · Store for tutorials33% (16/49)30% (18/61)43% (15/35)Task lightly reworded in Iteration 2. Wrong finishes cluster on a plausible sibling category.
9 · Add product video78% (38/49)52% (32/61)71% (25/35)Iteration 1 removed one correct destination, so the drop is partly a scoring change; excluded from improvement claims.

Percent Direct per task; n varies by task. All Iteration 2 figures are directional (38 sessions, several tasks reworded).

Sources

Study data: UXbeam tree-testing sessions and version history shown on this page.

Tree testing built for iteration

Iterate in one click. Compare versions. No participant caps. Full analytics included.