When the answer is the easy part
Designing human judgment into an AI tool for a profession where conclusions have to be defended.
CLIENT
Materia
Timeline
6 months
Role
Product designer (part-time)
The product had a magic button that would generate every answer at once.
The promise: you import a list of criteria based on what your audit team reads contracts against, and Materia's AI evaluates compliance for you, hundreds of times faster than human reviewers ever could.
This was 2024; the engine behind the button was still being figured out, and the quality of its output was an open question while I was designing its UX. That's why the stated design problem of making those answers good was beside the point for me. I couldn't design quality into an output that didn't exist yet.
What mattered was after. An auditor won't use a judgment they didn't derive themselves, regardless of how the answer turned out, because they're liable for the conclusions they draw.
So the design problem was what the human does after the button. And the human never blindly accepts.
I was brought in to design the review experience. What the work became was two reframes about what human oversight actually requires, plus a third I could see clearly but couldn't finish inside the scope.
Reframe 1: From correcting the output to working on it
The first build treated the human as a corrector. The AI produced an answer to be edited in the detail panel; the job was to fix what was wrong and move on.
That model is wrong about the work. Correction assumes the AI is mostly right and the human is a quality gate. But technical accounting judgment isn't gatekeeping. The human is doing the analysis, with the AI as a research assistant that got there first. The panel had to be a place where work happened, not where output got rubber-stamped. Reframing it from edit form to work surface changed what it had to hold.
One tradeoff was deliberate. In the audit context, a preparer needs to easily move between related items. So I prioritized traversability over immersion, and the panel kept a navigator for moving between line items instead of becoming a single-item, full-screen immersive experience. The table was always one click away.
Reframe 2: From a meta-assistant to context when it matters
A work surface for an AI product needs AI assistance inside it. The question was where it lived.
The team's first model was a meta-chat: one assistant, aware of the whole project, that you'd consult alongside the review. I argued for the opposite — a chat scoped to the individual line item, preloaded with that question's criteria, judgment, sources, and dependencies.
The team challenged it. Project-level interrogation is the more powerful-sounding feature; why give it up? My argument was a forced tradeoff: go project-level and you sacrifice either precision or speed, and in this domain neither is available to sacrifice.
Precision is a no-go because of the domain. A conversation that can't see exactly what you're judging can't help you decide whether this judgment, against these sources, is one you'll sign. Accounting doesn't tolerate approximate context.
Speed is a no-go because of the demo. A scoped chat opens already loaded; the preparer never re-orients the AI. That immediacy is part of the "magic button" moment, the thing that makes the product land in a demo. A meta-chat reintroduces setup cost at exactly the wrong moment.
Then the part that actually settled it: dependency tracing already solved the real version of the problem. Scoped chats can reference the items a judgment depends on, so cross-item context arrives when it's relevant instead of all the time. You don't need global context. You need context when it matters.
That reasoning is what shipped. The team killed project-level interrogation because the precision/speed tradeoff held and the dependency path made the global version unnecessary. The unit of judgment in this domain is the line item, and the assistant had to meet it there.
The scoped-chat decision also led to an idea inspired by MacOS Spotlight Search: an omni-chat, initiated anywhere, able to be pointed at any document, project, or thread in Conversations (another product feature). I wasn't there to build it past a knowledge-base lookup, but my reframe wrote the roadmap: once context when it matters replaces context all the time, a global chat router became the obvious next move.
What shipped
Due to confidentiality restrictions, I’ve omitted proprietary product screens. Wireframes and user flows are available upon request.
The thinking landed in two surfaces:
The table — the audit overview: status, flags, and filtering across every line item. This was the top-level review surface. The two-role thinking that would complete it was out of scope (see What I'd push further).
The detail panel — the preparer's work surface. Sources as verification infrastructure, dependencies as cross-item context, activity as an accountability trail. Includes the scoped "Chat with Materia" for line-item discussion, preloaded with the judgment's full context, sources, and dependencies.
Outcomes
The customer's worry was that a project-wide chat would pull in irrelevant context and degrade its answers. Scoping the chat to the line item removes that failure mode by design, not by tuning. That principle, scoped context over global context, shaped the roadmap for AI assistance after it.
Materia was acquired by Thomson Reuters in October 2024. The architecture the review experience sat on was sound enough to carry into a much larger product portfolio — a weaker, less attributable signal than the one above, but a real one about the durability of the underlying model.
What I'd push further
The biggest unfinished piece is a reframe the build never caught up to. By the end, the detail panel had become a sophisticated preparer's workspace, which raised a question the product hadn't answered: what does the reviewer need?
I recognized its shape from the pharma industry, where all designs are evaluated by MLR reviewers, the people who approve every claim before they go out, who don't write the material but own the sign-off and carry the liability. Audit depends on the same role, and it's not the same person as the preparer.
What they do | Where they live | What they need | |
|---|---|---|---|
Preparer | Derives and owns each judgment | The detail panel | A work surface: sources, dependencies, scoped interrogation |
Reviewer | Scans, samples, signs off | The table | A triage surface: flags, status at a glance, a path to the items that matter |
The table shipped as that surface, but only as a first pass — an overview, not a true reviewer's instrument. Given time, I'd build three things: a summary layer as the reviewer's entry point, sampling infrastructure as a principled shortcut to the items that warrant the most attention, and a sign-off mechanism as a way to transfer accountability, serving a different mental model from tracking progress.
There's a longer arc beyond this product. When Materia was built, the unit of judgment was the line item and the human had to traverse the audit by hand. As context windows have grown, so has the feasibility of larger engagements, but the human still has to check everything manually. The paradigm shift of human-in-the-loop UX is a system that watches the thinking continuously, notices what changed, and surfaces only what needs a decision, so the human's job is to judge, not to hunt for what needs judging.
It's a different product with the same starting point: judgment belongs to the human, and the system's job is to make that judgment as informed, and as frictionless, as it can be. It's what I'm building with Ariana.