Platform

Tree Testing and Card Sorting With Optimal Workshop

Test the structure before the design.

Optimal Workshop is where a proposed information architecture is put in front of the people who have to use it — tree tests, card sorts and first-click studies, with participants recruited inside the same tool.

Where it stands
In production

We run tree tests and card sorts in Optimal Workshop across museum, hospital, association and public health rebuilds — five of them written up here, each one validating the structure before any visual design existed.

Optimal Workshop is the research platform for testing a structure before it ships: tree tests, card sorts, first-click studies and a recruitment panel of more than ten million participants. Navigation is usually validated by whether the people who commissioned it like the wireframe, which puts the design and the politics on trial and leaves the structure untested. A tree test removes everything but the menu and the words in it: no layout, no color, no search box, a list and a task. What comes back is a findability rate for every task and the path each participant actually took, including the confident wrong turns — which are the ones a stakeholder review can never surface, because everyone in that room already knows where things live. The order is what makes it worth doing. Test the structure before anything is drawn and the architecture is the thing being judged; test it afterwards and the structure has already been reverse-engineered from a layout someone has approved, and the findings arrive as a list of changes nobody has the budget to make.

How the Work Splits

Optimal Workshop provides

Tree testing with per-task findability and the paths participants took; card sorting for how an audience groups and names things; first-click and prototype testing; surveys, interview analysis and live-site studies; a recruitment panel of more than ten million participants across 150 countries. SOC 2, GDPR, ISO/IEC 27001 and 27701, with WCAG 2.1 conformance in the study interface itself — which matters when the audience being tested includes people using assistive technology.

Pare & Co provides

The study design, which is where a tree test is won or lost. Tasks written in the audience’s words rather than the organization’s; a candidate structure worth testing rather than the current one with the labels changed; recruitment that reaches the people who actually use the thing; the reading of the result, where a 60% findability rate on one task and 61% on another routinely mean two completely different problems. Then the part no tool does at all — turning the paths people took into a structure and defending that structure to the departments whose pages moved.

Together

An information architecture the audience has already argued with. On the Isabella Stewart Gardner Museum rebuild the navigation and menu hierarchy were rebuilt structurally and tested before any visual design, against journeys mapped for the four audiences a museum site serves — members, first-time ticket buyers, on-site visitors and online researchers — so the structure was not reverse-engineered from a layout.

The work in practice

Most site structures are decided in a room. Someone proposes a menu, the people who know the organization best read it, everyone agrees it is clearer than what is there now and it ships. It is clearer than what is there now. It is also written by people who could find anything on that site blindfolded, which is the one qualification the audience does not have.

What a tree test actually measures

A tree test strips a proposed structure back to a list of labels and asks a participant to complete a task inside it. There is no layout, no visual hierarchy and no search box to rescue a bad menu — so the only things being judged are the shape of the tree and the words on it.

Two numbers come back and the second is the useful one. Findability is the share of participants who landed in the right place, and on its own it says a task went badly without saying why. The paths say why: whether people went to the wrong branch and stayed, went to the wrong branch and recovered, or wandered the top level because nothing there sounded like what they were looking for. Those are three different failures with three different fixes, and only one of them is solved by renaming something.

Directness matters as much as success. A task with a decent success rate and a lot of backtracking is a menu people can eventually beat rather than one they can read, and it is the pattern most likely to survive a stakeholder review unnoticed.

Card sorting is the other half, and it comes first

A tree test judges a structure you have. A card sort produces one you do not — participants group and name the content themselves, and the groupings that recur across enough of them are a structure that came from the audience rather than the org chart.

The pairing is the method: sort to generate the tree, tree test to prove it and iterate the labels that failed. Running only the second half means testing a structure someone invented, which is a faster way to find out it is wrong and no way at all to find out what would be right.

What the tool does not do

Recruitment, task wording and interpretation are the study, and all three sit outside the software.

Tasks have to be written in the audience’s words, and they may not contain the label that answers them — a task saying “find the membership renewal page” tests reading comprehension rather than the structure. The task is the situation the person is in: they joined last year and want to keep going.

The panel has to match the audience. A general panel will complete a public health task and produce a findability rate that describes the general public rather than the people in crisis the site is for.

And the result has to be turned into a decision. A failed task is a finding; moving a section is a change that costs a department its front door. The work after the study is the argument, and it goes better when the paths are on the table.

Where we have run it

On the Isabella Stewart Gardner Museum rebuild, navigation and menu hierarchy were rebuilt structurally and tested before any visual design existed, against journeys mapped for members, first-time ticket buyers, on-site visitors and online researchers. The order was the point: design first and validate afterwards, and the structure ends up reverse-engineered from the layout.

The same order runs through the rest of the IA work in the portfolio — a two-door architecture for a problem gambling helpline put through a twenty-five participant tree test of six tasks before design began, a four-site hospital consolidation whose patient-language navigation was validated with more than twenty-five participants and a member platform whose tree came out of a two-hundred-card sort run with six groups from across the organization.

If you are weighing it

The question is not whether the new navigation is better than the old one. It almost always is, and that comparison is why so many bad structures ship. The question is whether the people who will use it can find things in it, and the only way to know that before launch costs a fortnight and a panel.

Practice leadership

Need the navigation tested before it ships?

Tell us about it.
Start a conversation