The strongest near-term AI opportunity in complex operations is not replacing the domain expert. It is giving that expert better leverage.
People in regulated and high-consequence workflows do more than move information. They interpret incomplete evidence, recognise unusual situations, balance competing objectives and remain accountable for the decision. Much of their time, however, is consumed by searching, preparing, documenting and coordinating.
That is the product opportunity: reduce the administrative surface around judgement while preserving the judgement itself.
Start with where expertise is scarce
The most valuable users are not always the largest group. They may be quality specialists, clinical operations professionals, investigators, engineers, legal reviewers, programme leads or customer-facing experts. Their decisions shape safety, delivery, trust and cost.
I would begin by observing their work rather than asking where they want a chatbot.
What evidence do they assemble repeatedly? Which hand-offs create delay? Where are cases reopened because the record is incomplete? Which decisions require comparison across policies, countries or systems? What information arrives too late?
These questions reveal leverage points that a model demonstration will miss.
Preparation is a strong first product
Complex work often begins with preparation. The expert needs a concise view of the case, relevant history, applicable guidance, missing information and unresolved questions.
An AI product can assemble this material from permitted sources and make provenance visible. It can identify contradictions without pretending to resolve them. It can show what is known, what is inferred and what is missing.
That can improve speed and consistency while keeping the expert in control. It is also easier to evaluate than an autonomous decision. Reviewers can compare the prepared file with the underlying evidence and measure omissions, unsupported statements and time saved.
The interface matters. Evidence should be one click away. Uncertainty should not be hidden by fluent language. The product should make correction easy and preserve the correction as feedback for future evaluation.
Follow-through is another leverage point
After a decision, experts often translate the same outcome into several forms: a case note, a task, a status update, a stakeholder message and a structured system record.
A bounded agent can prepare those artefacts, route them to the right owner and track completion. The expert approves the substance; the system reduces duplication and missed follow-up.
The boundary must be explicit. Drafting is different from sending. Preparing a record is different from committing it. Suggesting the next action is different from executing an irreversible step.
This prepare-review-execute pattern is useful because the control is part of the workflow, not a separate governance meeting.
Build the product around evidence
Generative systems can produce plausible text from incomplete or irrelevant context. Domain experts need more than a polished answer.
I would design for:
- source-level traceability;
- clear separation of records, approved guidance and inference;
- date and version visibility;
- explicit missing or conflicting evidence;
- role-based access to sensitive information;
- a meaningful human approval step;
- a record of edits, overrides and final outcomes.
NIST’s Generative AI Profile, published in July 2024 and updated in April 2026, frames risk management across the AI lifecycle. That is particularly relevant here: quality is not only a model property. It depends on data, context, interface, human review and operations.
Do not automate ambiguity away
Some operational variation is waste. Some is the signal that expert attention is required.
An agent that forces every case through the same path may improve average speed while making unusual cases less visible. A good product distinguishes routine variation from consequential ambiguity.
That can mean confidence thresholds, explicit escalation reasons, exception queues and a safe way to stop. It can also mean showing two plausible interpretations rather than selecting one.
The product team should test cases where the correct behaviour is to ask for more information, defer to a specialist or do nothing. Refusal and escalation are part of quality.
Adoption depends on professional trust
Experts are often described as resistant to change when they are correctly sceptical of a product that hides its evidence or does not fit their responsibilities.
Trust grows when the product solves a real burden, respects the user’s role and responds visibly to feedback. It weakens when adoption is measured only by logins or when users must correct the same failure repeatedly.
I would involve representative users early, including people who handle edge cases. Training should use real workflow scenarios, not only feature tours. Local champions can support colleagues, but the product team still owns usability, reliability and the feedback loop.
BCG’s “From Potential to Profit: Closing the AI Impact Gap” (15 January 2025) places substantial weight on people, process and organisational change. That is not separate from the product. It is how the product reaches an outcome.
Measure leverage, not displacement
The most useful measures are close to the work.
How long does preparation take? How often is evidence missing? How many cases are reopened? How much editing is required? Are experts able to handle more work without reducing quality? Do users report better decision confidence? Are exceptions identified earlier?
Headcount reduction is a poor default objective for a product that supports scarce expertise. The better aim is to increase capacity, consistency, responsiveness and learning while keeping accountability clear.
McKinsey’s 2026 global AI survey reports broad individual productivity benefits but a persistent gap to enterprise-level impact. Moving from personal assistance to redesigned operational workflows is where product leadership becomes decisive.
Keep the human role explicit
“Human in the loop” can become a slogan. The design needs to specify what the human sees, decides and owns.
For low-consequence work, review may be sampled. For sensitive communication, a named person may approve every output. For a regulated or safety-relevant decision, the system may only prepare evidence while a qualified professional determines the outcome.
As performance improves, the boundary can change—but only with evidence. The purpose is not to preserve manual work indefinitely. It is to automate where the system is dependable and keep human judgement where context, accountability or consequence requires it.
Three takeaways
- Look for administrative work surrounding scarce judgement: preparation, evidence assembly, documentation and follow-through.
- Design for provenance, exceptions and meaningful approval rather than fluent output alone.
- Measure capacity, quality and responsiveness; treat professional trust and adoption as product outcomes.
*The views expressed here are personal and do not represent those of my current or previous employers.*