Information architecture / Assessment

Information architecture assessment

An information architecture assessment identifies likely structural or findability issues, the evidence already available, and what to change or test next. It ends with an actionable decision and the evidence supporting it.

It is not a whole-interface usability test, a content audit, or a redesign, though it may show that one is needed.

Deliverables: a findings table by structural layer, task-level results where tested, and a decision with a validation plan.

Before recruiting participants

  1. Write down the decision in one sentence: what is being decided and what is allowed to change.
  2. Pull your site-search queries. Start with the most frequent queries and the ones that returned no results. They reveal search vocabulary and recurring content needs.
  3. List the priority tasks and the audiences that perform them. Use them to decide what matters most to inspect or test.

Three sizes of assessment

Size depends on what you need to claim, not on the size of the site.

SizeWhat it involvesWhat it can showWhat it cannot show
Evidence onlySteps 1 to 3 below, using material you already haveWhich part of the structure is the likely problem; gaps in coverage and vocabulary; contradictions between sourcesFindability
One targeted studySteps 1 to 3, plus one tree test of the current structure, or a card sort if grouping is the open questionTree test: how findable the tested tasks are in the structure. Card sort: candidate groupingsHow the live interface affects finding; tasks you did not test
Comparison or fullAdds a like-for-like comparison of structures, and/or usability testing of the live interfaceHow structures perform relative to each other on the same tasks; effects of the interfaceCauses outside what you tested

From the worked example: two candidate structures were tree-tested on eight tasks, and one appeared about 10 points higher overall. Three of the eight tasks had been worded differently for each test. On the five worded identically, the structures were close. Read the example

Is it an IA problem, and which part?

Each row identifies the likely part of the structure and the first diagnostic check. A check is not a fix.

What you seeLikely partFirst checkWhere evidence may already exist
Frequent searches for content already in navigationLabels, or hierarchy and depthCompare query terms with route labelsSite-search logs
Findability complaints about content that existsLabels, or hierarchy and depthEntry pages and required navigation pathSupport contacts; analytics entry pages
Zero-result searches for content you do not offerCoverageConfirm the content is really absent, not just named differentlySite-search logs; the content owner
The same content exists in several placesCoverageWhich version receives traffic, and which one is currentContent inventory; analytics
Teams disagree about where something belongsGroupingExpected placement in one or multiple categoriesEarlier card sorts; support vocabulary
Correct section, wrong final destinationLabelsWhich neighbouring item they pick instead, and how its label comparesEarlier tree tests; analytics paths
Some content is reachable only through search or a direct linkHierarchy and depthWhether those tasks require a browse path to the contentInventory (pages with no links to them); analytics referrers
Users miss an item that is present in the live navigationOutside the structure: interface and navigation designFindability without interface cuesA tree test of the same task; compare with usability sessions on the live interface
Search finds the page but ranks it low, or returns irrelevant resultsOutside the structure: the search systemWhat the top queries actually returnSite-search logs
Users reach the page, but the content does not resolve the taskOutside the structure: content qualityTask-content fitSupport contacts; page feedback
Labels and placement drift as different teams add pagesOutside the structure: governanceWho decides where new content goesPublishing records

The four parts of an information architecture this page uses

UXbeam uses this four-part synthesis of established practice to organise findings. It is not a standard model.

PartIts questionTypical evidence
CoverageDoes the structure cover content required for priority tasks?Zero-result searches, support themes, an inventory where warranted 56
GroupingDo groupings match user expectations?Card sorts; incorrect tree-test destinations 47
LabelsDo labels support correct first clicks and destinations?Search vocabulary; first clicks 16
Hierarchy and depthAre target destinations findable through the hierarchy alone?Tree tests 128
  • Grouping is not findability. Card sorting reveals participant grouping patterns; it does not establish navigation findability 14.
  • The structure is not the interface. The same structure can perform differently depending on how the interface presents it. Test structure and interface separately when you need to isolate the cause 3.
  • Two studies are not automatically a comparison. The conditions are under Reading results.

A tree test measures findability: success in locating an expected destination through the structure 3.

Run the assessment in five steps

Three steps include a STOP condition. Stopping early is valid when the decision is already resolved or nothing can change.

Copied

1. Define the decision and what can change

Write the decision in one sentence, for example "keep or replace the top-level menu before the spring release". Then list what is fixed (platform limits, legal names, deadlines), what can change, and who decides. This becomes the scope: you only need evidence about the parts that can change.

Questions to answer before going further:

  • Which audiences matter most?
  • What would make the decision obvious?

Stop here if nothing can change. More evidence will not change the decision.

2. Gather what you already have

Use existing material before you recruit anyone. Most of it sits with other teams, so ask early. Rows marked "useful" are optional; "conditional" means only in specific cases.

MaterialStatusWho usually has itMinimum useful amountUse it toDon't use it to
The decision and what can changeRequiredSponsor or product ownerOne sentence plus constraintsSet the scopeTreat it as evidence
The structure: the current tree, the draft if you are comparing, tags if you use themRequiredCMS or product ownerEvery level you will testSee shape, labels and depthAssess findability 1
Top tasks and who does themRequired for any test with participantsProduct, support, analyticsThe tasks that matter most, by audienceChoose tasks and participantsCover the long tail
Site-search logsUseful; start here if the site has searchSearch or analytics ownerThe top 1,000 queries over 3 to 6 months; work through the first 300 or so 6Identify search vocabulary and possible content gapsExplain search intent or represent non-searchers 10
Support and contact themesUsefulSupport leadRecent contacts, tagged by themeCapture support language and reported problemsEstimate prevalence
AnalyticsUsefulAnalytics ownerTop entry pages and most-used pagesIdentify high-use content and entry pagesInfer intent or causation 9
Earlier research, including any baselineUseful; required to claim improvementResearch repositoryDate, tasks, number of participants, test formatReuse tasks and starting questionsCompare with new results, unless tasks, wording, scoring and test format match
Content inventoryConditionalCMS ownerA sample, or one section, within a fixed timeCheck coverage, duplication, outdated content, and pages with no inbound linksAssess findability 5

Build an inventory only when coverage or duplication is the suspected problem, or when the structure cannot be listed another way, and only after Step 1. NN/g gives six weeks as one example for a first pass 5.

Reuse earlier tree-test tasks when they are still comparable. They let you re-test like for like. Tasks you write now become the baseline for the next round.

Participants, if you will need them

  • Who: participants from each audience the decision affects. Results from one audience do not generalize to another.
  • Where from: your own users, a recruitment panel, or colleagues. Colleagues are fine for a pilot to check the tasks, but not for study results.
  • How many depends on the purpose. A qualitative pilot needs a few participants 1. For quantitative tree testing, about 30 participants seeing each task can start to show patterns within an audience, and about 50 seeing each task is a common target 11. If each participant sees only part of the task bank, recruit proportionately more participants so each task still gets that coverage. Comparing two structures typically needs about 50 or more observations per task, per structure for reasonably narrow confidence intervals 2. Each step up adds recruitment cost and time.

Review the structure itself

You can do this today with no participants: open the sitemap and inspect the structure. None of the conditions below is a defect on its own; the table also shows when each may be acceptable. A desk review starts the assessment; it is not the final word 12.

ConditionWhy lookWhen it may be fine
The same page reachable by several routesMultiple routes split behavior and make measurement ambiguous unless the route is knownOften intentional and helpful, especially where audiences differ
Content reachable only by search or a direct linkNo browse path reaches the contentSometimes correct by design: deep reference material, archives, or content reached from external links
A branch with a single childA level that offers no choice costs a stepIt can give useful context or a stable address, or leave room to grow
Siblings of different kinds at one levelThe grouping principle may be inconsistent at this levelA prominent task next to topic sections can be deliberate
One section holds a large share of the itemsThe grouping may have become too broadUneven section sizes may reflect the domain
Items with more than one plausible homeA good source of tasks to testTwo plausible homes are normal; test expected placement
Important content placed deeplyDeep placement adds navigation effort to important contentDepth alone does not decide findability

Rules that mislead when applied everywhere

  • Three clicks. What matters is whether each choice is predictable. A clear five-step path can work better than an ambiguous two-step one.
  • Seven plus or minus two per level. This comes from working-memory research, not studies of menu scanning.
  • No more than N levels. There is no universal correct depth.
  • A structure health score. Tree depth and shape do not establish findability.

3. Read the evidence and identify the likely part

Go back to the triage table with what you gathered and file each signal under one of the four parts, or outside the structure. Then look for places where sources disagree. Check whether they cover the same task, audience and setting before treating either as decisive.

Contradictions to look for

  • Search logs show frequent queries for content already in navigation. The label may not match search vocabulary, the route may be hard to follow, or search may simply be preferred.
  • Support reports findability failures while analytics shows high usage. The support audience may differ or arrive by a different route.
  • An earlier study found no problem, but complaints continue. Check whether it covered the complaint themes with the same audience.
  • The team agrees on placement, but search vocabulary differs from the label. The grouping may be sound while the label is not.

Reading search words. Compare query terms used for the same need and the audiences using them. Search logs often contain specific-item terms rather than category labels, so raw query frequency is not a direct measure of label quality. A frequent query matching the label shows label recognition, not browse findability.

What counts as a complaint. Classify a support request as a findability complaint only when it describes a failed attempt to locate content or complete a task.

Coverage gaps in existing evidence. Analytics, search logs and support records cover only observed visits, searches and submitted contacts. If a target audience is absent from those sources, you still lack findability evidence for that audience.

Stop here if the evidence already resolves the decision: no plausible study result would change it, and the affected audience is represented in the evidence. Do not use this stop for critical or costly-to-reverse changes, and do not report the structure as tested.

4. Test only what remains open

Define the unresolved research question in one sentence, then select the method that answers it.

MethodAnswersCannot tell youDo not use it whenTypical sizeUXbeam
Desk or expert reviewStructural and labeling issues identified by expert reviewTask success 12It would be the only evidence for a consequential changeThree or more independent reviewers if you rate severity 13
Open card sortParticipant-generated groupings and category labelsNavigation findability; task performance 14The categories are fixed and cannot changeAt least 15 participants for a qualitative sort; 30 to 50 or more for a quantitative one; 30 to 50 cards 4✓
Closed or hybrid card sortFit of items within supplied categoriesNavigation findability; covers one level 314The goal is findability; use a tree test instead 4Use the open-sort guidance above 4✓
Tree testFindability, first clicks and navigation paths; comparison between structuresEffects of the interface, search, related links or page content 2The interface is the suspected causeA few participants for a qualitative pilot 1; about 30 participants seeing each task for early patterns, about 50 as a common target 11; recruit proportionately more if participants see only part of the task bank; comparisons about 50 or more observations per task, per structure 2; 8 to 10 tasks per participant 15✓
First-click testFirst-click choice on a screen or mockupBehavior after the first click 3The navigation depends on hover menus, which a static image does not reproduce well 16Not covered here
Usability test (prototype or live)Task performance with structure, interface and content combinedWhich component caused a failure 14You need to isolate the structure from interface and content effectsNot covered here

UXbeam's recommended default order: existing evidence first; then a baseline of the current structure (a tree test isolates the hierarchy; a usability test covers the whole experience) 1817; a card sort when grouping is the open question or a new structure is being built; moderated sessions to explain a failure you have already located. NN/g describes a common alternative: usability test, then card sort, then tree test 1.

Change the order when the situation calls for it: if you already have a draft and no current structure to baseline, tree-test the draft directly. If the grouping is still open, card sort first. A fixed set of categories makes an open sort the wrong first step; a very small site may need no tree test at all.

Tasks

Take tasks from the top-task list, search logs and support themes. Give each participant 8 to 10 tasks; more can create learning effects 15. Write tasks in user language and avoid tree-label wording; otherwise the task becomes a word-matching exercise 9.

Look for semantic cues as well as exact words. "Your green bin wasn't emptied" can cue the Environment section without sharing its label. When a frequent query matches the label, write the task around the underlying need ("you want to know what a mortgage would cost you each month"), not the query.

More on task design: writing tree test tasks and avoiding leading tasks.

Comparing a current and a proposed structure

Test the current structure as a baseline. Use the same tasks, wording, scoring and test format for both, with comparable participant populations. See Reading results for the limits of that comparison.

What drives effort: the number of audiences, whether you compare structures, task allocation per participant, the size of the tree you test (not the size of the site), and where participants come from.

A small qualitative study can run with a spreadsheet and a form. You do not need a dedicated tool to start.

Stop when no method answers the remaining question. Report it as open, with what would be needed to answer it.

Need to test findability?

Start a tree test

Need to test content grouping?

Start a card sort

5. Read the results and decide

Tree tests report three outcomes per attempt: direct (reached a correct item without backtracking), indirect (reached it after backtracking) and elsewhere (finished at an item not marked correct). Success means direct plus indirect. Read success, directness, first clicks and final destinations together; a rate alone does not tell you what to change 2.

Patterns to notice

Each pattern supports an interpretation, not a verdict. The last column shows what to inspect before changing anything.

What you seeWhat it might suggestWhat to inspect next
On a low-scoring task, wrong first clicks cluster on one other sectionThat section may match the expected location, or task wording may cue itFinal destinations within that section; task-label word or semantic overlap 2
First clicks are right, but most finish at one or two wrong items in the sectionThe labels at that level may overlap or be near-twinsThe list of final destinations, and how the attracting item's label compares with the correct one 2
First clicks scatter across several sectionsFor one or two tasks, the item may have several plausible homes. For many tasks, the top level may be unclearNumber of tasks affected. For one or two, use moderated think-aloud sessions 8. For many, run an open card sort 2.
Tree-test tasks score well, but findability complaints continueThe test may not have covered the reported problems, or the cause may sit outside the structureEach complaint theme against the tasks you tested; write and test tasks for the themes you missed 23
Repeated zero-result searches for content you do not offerA content-coverage issue, not a structural oneGroup the missing requests by theme and volume, then take them to the content owner 6
Strong card-sort agreement is cited as proof that the menu worksGrouping agreement does not establish findabilityCheck whether shared wording or a leftovers group explains the apparent agreement; then tree-test the proposed structure 1418
The tree test and a test of the live site or prototype disagreeThe two settings expose different routes and cuesObserved routes in each setting, and whether the live route exists in the tested tree 2319
Earlier research used different tasks, wording, scoring, test format or audienceUse it to generate questions and tasks, not as a baselineUnchanged tasks and branches; then apply the comparison conditions below 219

Possible changes, once an inspection supports them: move an item, list it in a second place, rename a section, make two neighbouring labels distinct, fix the content, or hand the problem to the interface or search owner. List an item in a second place only when a substantial share of participants choose that location; doing it everywhere dilutes the structure 2.

What results cannot tell you on their own

  • Comparable conditions. Two results compare only when the tasks and their wording, the scoring and the test format match, and the participant populations are comparable. Otherwise compare only the parts that match, and say so. A 2025 peer-reviewed study found that different tree-test formats gave significantly different results 19.
  • Modest samples leave moderate differences uncertain. This applies to success rates, to first-click shares and to final-destination shares. A gap between two results is less certain than either result on its own, so do not judge a gap by one result's margin of error.
  • "Not significant" does not mean "the same"; it means the study could not resolve the difference. A lower score on a critical task therefore remains unresolved, not a tie. A sample of 50 or more participants per structure, often recommended for comparisons, gives reasonably narrow intervals 2, but can still miss a moderate loss on a single task. If such a loss would matter, plan the study to detect it.
  • A re-test is evidence, not proof. A better score after a change fits the idea that the change helped. Before saying so, check what else changed (task wording, scoring, test format, audience, date, a new label that repeats a task word), whether the pattern that prompted the change moved, and the other tasks the change could affect.
  • Changing the answer key after seeing results. If a "wrong" destination looks acceptable, judge it on content alone: would that destination satisfy the task? If you change what counts as correct, report the result both ways. Keep the original scoring for comparison with earlier rounds; treat the rescored result as a different analysis.
  • No universal pass mark. There is no sourced universal threshold for a "good" structure. Interpret each task in light of its criticality and audience.

When the evidence points elsewhere

  • Content: required content is missing, inaccurate, or duplicated. It goes to the content owner.
  • Interface and navigation design: the item is findable in the tree but not on the live site. Test the interface 3.
  • Search: relevant content exists, but site search returns poor results. Route it to the search owner 6.
  • Grouping: first clicks scatter across many tasks, or placement remains disputed. Run a card sort (card sorting, reading card sort results, organizing principles).

Prioritize by task criticality and frequency. If you rate severity, average three or more raters 13. For content, choose keep, update or remove 5. Little published evidence covers prioritising IA findings specifically.

Worked example

Two structures for a connected-car app, and a comparison that was only partly like for like

The question. Companion apps for cars usually group features by vehicle system: climate, security, charging. Should a new app follow that convention, or organise features by whether you are near the car or away from it?

Evidence already available. An audit of existing companion apps gave a list of 24 features. None of the audited apps organised features by whether you are near the car.

Card sort. In an open card sort, the first 60 US adult participants sorted the 24 features. Most grouped them by vehicle system. About a quarter grouped them by a different rule: near the car, or away from it. The sort produced two plausible grouping schemes; it could not tell which would be easier to navigate.

Two candidate structures. The team built one structure on each organizing principle: A followed the vehicle-system convention, B the near-or-away principle. Both held all 24 features.

Tree test. Each structure was tested with a separate group of participants on eight tasks with the same intent, such as starting the car remotely or finding where it is parked. After excluding internal sessions, one outlier, and one participant who completed both tests, 18 participants per structure remained, with 144 attempts each.

Result. Overall success was 60% for A (87 of 144 attempts) and 70% for B (101 of 144). Direct success was identical: 63 of 144 for each.

What the comparison could not separate. Three of the eight tasks were worded differently in the two tests. In A, one read "make sure it's protected while you're gone"; in B, "turn on the security monitoring". Another read "make sure you have enough range" in A and "check your fuel level" in B, where B's correct item was labelled Fuel Level. These three tasks include the two largest gaps between the structures, and B's lead overall comes from them: 44 of 54 successes for B against 26 of 54 for A.

On the five tasks worded identically, the structures were close: 61 of 90 successes for A, 57 of 90 for B, and 42 of 90 direct for each. The clearest matched difference was directness on one task, unlocking the car from home: 3 of 18 went straight to the answer in A, 11 of 18 in B, with similar overall success.

Where the overall gap came from

Matched: A 61/90 · B 57/90. Mismatched: A 26/54 · B 44/54. All tasks: A 87/144 · B 101/144

Where the overall gap came from
TasksA · Vehicle systemB · Near or away
5 tasks worded the same for both61/9057/90
3 tasks worded differently26/5444/54
All tasks87/144101/144
Successful attempts per structure, 18 participants each. The overall difference comes from the three tasks whose wording differed between the two tests.

What it supports. Matched-task results did not favor one structure overall. B had higher Direct success on one remote task. Both remained plausible candidates.

What it does not support. The results do not support the claim that B is better overall: the overall gap comes from tasks that were not worded the same. With 18 participants per structure, differences of this size on single tasks remain uncertain. Tree testing does not measure live-interface performance.

What would settle it. Re-test the three mismatched tasks with one wording for both structures, written around the need rather than either label, and with more participants per structure if the choice matters.

A second case: an e-commerce admin, three rounds

  • The overall average hid large task-level changes. Between the first two rounds, on the same nine tasks (three lightly edited), overall direct success moved 3 points (30% to 33%). Underneath, two tasks gained 17 to 19 points and another lost 12.
  • In one menu, 26 of 54 participants (48%) chose "Auto send review request" instead of the correct near-twin label, "Auto Sent review requests".
  • The third round reached 48% direct after both the tree and several task instructions changed, so the gain cannot be attributed to the tree alone. Earlier, a section renamed "Customer Notifications" took on the word "notifications" from its own task, introducing a word-match confound 9.

The e-commerce case study

In one published government project, the original structure scored 31% success and the new structure 67% on the same tasks 8: one project, not a typical gain, and a baseline is only one of the comparison conditions above.

What you hand over

Minimum: the findings table, results per task where you tested, and the decision with its re-test plan.

Full: the minimum, plus an inventory and audit if coverage was in question 5, a revised structure tested again 9, and a short note on the limits of what was tested.

For stakeholders: lead with the decision, then the two or three findings that drive it, then what remains open and how it will be checked. Keep the full table as the appendix.

Findings table

Part of the structureData / observationInterpretationConfidence / evidence basisSeverity (if rated)OwnerNext step
Labels26 of 54 finished at "Auto send review request"; the correct item was "Auto Sent review requests", in the same menuThe two labels are near-twins26 of 54 chose the same wrong destination; one clear competing labelNot ratedProduct team for that menuRename one label; check both destinations in the next round

Keep data/observations and interpretations in separate columns.

What an assessment cannot tell you

  • Method-specific limits are stated where they apply: live-interface effects in Step 4, audience coverage in Step 3, uncertainty in Reading results, and expert judgment in the desk review.
  • Whether the structure will stay coherent over time. Study findings reflect the tested tasks, participant sample and test format at that time. Structures drift as content is added; ongoing governance must define where new content goes.

Teams that track top tasks over time re-measure them every 6 to 12 months; this is top-task measurement, a separate practice from tree testing 20.

Start with what you already have

Steps 1 to 3 can be done without a dedicated research platform or new participant recruitment. Go to step 2

Test findability:

Start a tree test

Test content grouping:

Start a card sort

Sources

  1. [1] Laubheimer, P. "Tree Testing: Fast, Iterative Evaluation of Menu Labels and Categories." Nielsen Norman Group, 2023.
  2. [2] Laubheimer, P. "Tree Testing Part 2: Interpreting the Results." Nielsen Norman Group, 2024.
  3. [3] Cardello, J. "Low Findability and Discoverability: Four Testing Methods to Identify the Causes." Nielsen Norman Group, 2014.
  4. [4] Tankala, S. and Sherwin, K. "Card Sorting: Uncover Users' Mental Models for Better Information Architecture." Nielsen Norman Group, 2024.
  5. [5] Kaley, A. "Content Inventory and Auditing 101." Nielsen Norman Group, 2020.
  6. [6] Farrell, S. "Site-Search Log Analysis." Nielsen Norman Group, 2017.
  7. [7] Spencer, D. and Warfel, T. "Card Sorting: A Definitive Guide." Boxes and Arrows, 2004.
  8. [8] O'Brien, D. "Tree Testing." Boxes and Arrows, 2009.
  9. [9] Spencer, D. A Practical Guide to Information Architecture, 2nd ed. UX Mastery, 2014.
  10. [10] Wiggins, R. and Rosenfeld, L. "Using Search Analytics to Diagnose What's Ailing Your Information Architecture." IA Summit, 2007.
  11. [11] O'Brien, D. "How many participants?" In Tree Testing for Websites.
  12. [12] Arango, J. "Mastering the Muddle: The Power of Heuristic Evaluations." 2023.
  13. [13] Nielsen, J. "Severity Ratings for Usability Problems." Nielsen Norman Group, 1994.
  14. [14] O'Brien, D. "Comparing to other IA methods." In Tree Testing for Websites.
  15. [15] O'Brien, D. "How many tasks?" In Tree Testing for Websites.
  16. [16] Sauro, J., Schiavone, W., Du, D. and Lewis, J. "Do Click Tests Predict Live-Site Clicks?" MeasuringU, 2023.
  17. [17] Canada.ca Experience Office. "How we're optimizing Canada.ca top tasks." 2017.
  18. [18] Nielsen, J. "Card Sorting: How Many Users to Test." Nielsen Norman Group, 2004.
  19. [19] Kuric, E., Demcak, P. and Krajcovic, M. "Validation of information architecture: cross-methodological comparison of tree testing variants and prototype user testing." Information and Software Technology 183, 2025. The authors work for a tree-testing vendor.
  20. [20] Wake, L. "Top task performance indicators." University of St Andrews Digital Communications, 2017 (reporting Gerry McGovern's method).

Study data: UXbeam card sorting and tree testing studies shown in the worked examples.