When the answer is the easy part

Designing human judgment into an AI tool for a profession where conclusions have to be defended.

CLIENT

Materia

Timeline

6 months

Role

Product designer (part-time)

The product had a magic button that would generate every answer at once.

The promise: you import a list of criteria based on what your audit team reads contracts against, and Materia's AI evaluates compliance for you, hundreds of times faster than human reviewers ever could.

This was 2024, and the engine behind the button was still being figured out. The quality of its output was an open question while I was designing its UX. That is why the stated design problem, making those answers good, was beside the point for me. I couldn't design quality into an output that didn't exist yet.

What mattered was after. An auditor won't use a judgment they didn't derive themselves, regardless of how the answer turned out, because they are liable for the conclusions they draw. So the design problem was what the human does after the button. And the human never blindly accepts.

I was brought in to design the review experience. The work became two moves. One was figuring out what reviewing an AI's judgment actually takes, since no one knew. The other was reframing what AI assistance is. Each left me with a principle I design by now, and Materia is where I worked them out.

Discovery: what reviewing an AI's judgment takes

No one on the team could say what reviewing an AI's judgment actually involved. The first build defaulted to an edit form, the AI's answer in a field to fix and move past, but that was a placeholder for an undefined problem, not a considered position. The work was to define it.

I built the user story map the team worked from. It laid out what reaching a judgment takes: read the AI's draft, check it against its sources, pull in the items it depends on, interrogate it, and only then commit to the conclusion you sign. That map turned a vague "review step" into a work surface, a place where the analysis happens rather than where output gets rubber-stamped.

The product's whole promise was speed, which creates pressure to make the human's part fast too. The honest way to do that is a strong first draft and the means to check it quickly, so the panel kept the evidence within reach: the sources a judgment rests on, the items it depends on, and the record of what was done. I prioritized traversability over immersion, since a preparer works across related items.

The principle it proved out: augment the judgment, don't shortcut it. The AI makes the first pass fast; the design makes the human's own pass fast and defensible.

Reframe: what AI assistance is

A work surface for an AI product needs assistance inside it, and the question was what that assistance should be. The team's instinct was that more powerful means more aware: a meta-chat that sees the whole project. I reframed it. What makes an assistant valuable is fit to the decision in front of you, not the breadth of what it can see. In this domain, fit beats reach.

That wasn't a guess. The product already had a separate workspace, Conversations, used mostly for research, and customer interviews had shown how people used the assistant there. Working a line item is essentially focused research, the same activity narrowed to one judgment and its criteria, sources, and dependencies. So the assistance should take the same form, a Conversation pointed at the line item, a scoped chat preloaded with exactly that context.

Making the case against the meta-chat was a forced tradeoff: going project-wide sacrifices precision, speed, or legibility, and in this domain none is available to give up.


Tradeoff

Scoped chat

Project-level meta-chat

Precision

Sees exactly what you're judging, against its own sources

Can't see the specific judgment, so context stays approximate, which accounting doesn't tolerate

Speed

Opens already loaded, so you never re-orient it

Reintroduces setup cost every time you consult it

Legibility

Keeps each item's thread with the item, so you can see how you worked it through

Trails a thread off every line item into one stream, a nightmare to navigate

Then the part that settled it: dependency tracing already solved the real version of the problem. A scoped chat can reference the items a judgment depends on, so cross-item context arrives when it is relevant instead of all the time. You don't need global context. You need context when it matters.

That reasoning shipped, with project-level interrogation deferred to the roadmap as a lower priority. It also set the shape the global version would take: once context when it matters replaces context all the time, that version is a chat router, an omni-chat you point at any document, project, or thread in Conversations, rather than a standing meta-chat.

The principle it left me with: scope the assistant to why people actually talk to it. Reach sounds powerful; fit is what makes it useful.

What shipped

Proprietary screens are omitted. The wireframes and flow below are abstracted for confidentiality: the real structure and interaction logic, with all content removed.

The table is the audit overview: status, flags, and filtering across every line item, and the way into each one. It shipped as a first-pass overview. A fuller reviewer's instrument was out of scope, the surface for the reviewer who scans, samples, and signs off, and carries the liability without deriving each judgment. I knew that role from pharma's medical, legal, and regulatory (MLR) review, and it is not the same person as the preparer.

The detail panel is the preparer's work surface: sources as verification infrastructure, dependencies as cross-item context, and activity as an accountability trail. It includes the scoped "Chat with Materia" for line-item discussion, preloaded with the judgment's full context, sources, and dependencies.

The signature interaction lives in that panel. The AI proposes an answer with its sources; the preparer verifies it against those sources, traces the dependencies for cross-item context, and interrogates the scoped chat; then they derive and own the judgment they sign. Agreeing, adjusting, or rejecting is the outcome of that work, not the interaction.

Outcomes

The customer's worry was that a project-wide chat would pull in irrelevant context and degrade its answers. Scoping the chat to the line item removes that failure mode by design, not by tuning. That principle, scoped context over global context, shaped the roadmap for AI assistance after it.

Materia was acquired by Thomson Reuters in October 2024. The architecture the review experience sat on was sound enough to carry into a much larger product portfolio, a weaker and less attributable signal than the one above, but a real one about the durability of the underlying model.

Applying a 2026 lens

The two principles still hold. What 2026 changes is where the first one lives.

Looking back from 2026, the 2024 approach concentrated the oversight after the run. Users imported the criteria to audit against, the AI evaluated against them, and the human reviewed each output. That held when a person could read every output. As agents take on more, reviewing everything after the fact is where it falls short.

The same shift is reaching MLR, a regulated review with the same shape as audit. It has been top of mind because in my own work we have been discussing processes and internal tools that use AI to facilitate that review, with agents doing the checking and the human keeping the sign-off.

A 2026 study of developers overseeing software agents, small and qualitative, finds oversight is "not only reactive and retrospective... but also preventative and proactive," and names four forms: a priori control, co-planning, real-time monitoring, and post hoc review ("Human oversight of agentic systems in practice," arXiv:2606.05391). The current approach uses only the last of the four; the shift is toward the first three, before and during the run.

In an audit product, that is the first principle relocated: the human still owns the judgment, now exercised in the plan they author before the run rather than output by output. Criteria still define what the agent checks against, and risk sets three conditions, what the agent may decide on its own, how much evidence an output must clear, and what must be escalated. Review doesn't go away. The auditor still signs and carries the liability, now reviewing what risk escalates, with a traceable basis. The concept below sketches it: the plan the user sets, and the review it produces.

There's a longer arc beyond this product. When Materia was built, the unit of judgment was the line item and the human had to traverse the audit by hand. As context windows have grown, so has the feasibility of larger engagements, but the human still has to check everything manually. The paradigm shift of human-in-the-loop UX is a system that watches the thinking continuously, notices what changed, and surfaces only what needs a decision, so the human's job is to judge, not to hunt for what needs judging.

It's a different product with the same starting point: judgment belongs to the human, and the system's job is to make that judgment as informed, and as frictionless, as it can be. It's what I'm building with Ariana.