Enable javascript in your browser for better experience. Need to know to enable it? Go here.

Beyond the quick-win: Designing AI systems around critical human judgment

Why quick wins fail, and how understanding the bigger picture protects your team, the output quality and your investment

In discussions around AI, the phrase "human in the loop" is often treated with fatigue. Yet in many enterprise applications, human oversight remains essential to delivering value safely and effectively. "Human in the loop" is often discussed as a temporary fix: something we have to live with until AI gets smart enough to do the whole job alone. Some practitioners view human involvement as something that can eventually be eliminated once enough data has been collected.

 

This is a mistake.

 

When AI projects fail to deliver real value, the problem may lie less with model capability than with the design of the wider workflow. Many solutions use models far more sophisticated and costly than necessary. It is often because the overall workflow was poorly designed. Teams have failed to take a wider view of the system.

 

Under pressure to show fast results, teams rush for quick wins. They build chatbots for routine customer support or quick tools to speed up basic paperwork. The problem is this is happening across the organization simultaneously, resulting in chaos. 

 

At first, the numbers look great: call volumes drop and short-term metrics jump. But if no one is looking at the whole system, the bottleneck may not disappear, it may just move somewhere else. 

 

Rather than chasing low-margin automation quick wins, business leaders should identify the systemic leverage points where human judgment is irreplaceable. These points require deep consideration rather than rapid patches, but they yield the highest long-term return on investment.

 

When AI takes away all the easy, routine tasks, human staff are left handling only the hardest, most frustrating edge cases. Without easy tasks to buffer their day, workers burn out. Resolution times go up, service quality drops, and the initial savings vanish.

 

Human judgment is not a cost to eliminate. It is the most valuable part of your system when things get complicated. 

 

 

The chatbot problem: A customer support example 

 

To see what happens in practice, it's worth looking at a standard customer support rollout. Many SaaS customer service platforms now include chatbots capabilities, making them relatively easy to deploy. 

 

A business deploys an AI chatbot to handle routine operations such as password reset, order tracking (“where is my order?” or WISMO), store opening times, etc. On paper the initiative feels like an instant win. Inbound call volumes drop, query resolution jumps and first week metrics look excellent on exec dashboards. 

 

Two months later, the systemic effects kick in.

 

Before the chatbot, a call centre agent’s day was a mix of easy tasks and more complex enquiries: it made the day interesting and gave them a feeling of satisfaction, they liked to solve customer problems. 

 

When the chatbot takes over all the easy work, the buffer and reward vanishes overnight. 

 

Suddenly, much of the work that hits human staff consists of edge cases:frustrated customers, complex problems or multi system challenges. 

 

The ‘humans in the loop’ hit a cognitive wall. 

 

  1. They spend all day making difficult decisions and judgements under pressure.

     

  2. Call handling times increase as complex enquiries require talking to other humans, checking multiple systems and digging through historical context.

     

  3. Service quality begins to degrade, tired agents make more mistakes and CSAT and NPS scores begin to drop.

     

 

Because the chatbot was only looking at initial problems, it didn’t fix the system, it pushed more of the complex work around to another point: the human in the loop. 

 

 

Design with the big picture: Feedback loops and system lag

 

It would be wrong to say fixing this means removing the AI from the picture. But perhaps we need to change how we design systems around AI. 

 

When you roll out an AI tool, the system effects rarely appear on day one. There is often a system lag: a delay between launching the quick-fix automation and seeing secondary impacts such as team burnout, customer retention or supplier contract margins. 

 

Without considering the whole system, leaders may interpret the problems as operational: poor performance, weak decision-making or supplier challenges. 

Mapping the system beyond the user journey 

 

Instead of drawing simple user journeys and isolating specific failure points, product and technology leaders can map and consider three sorts of feedback loops. 

 

  1. Efficiency loops: Where will the immediate savings come, and how much needs to be reinvested in system changes?

     

  2. Fatigue loops: How do we design the work reaching human agents so that it remains manageable and rewarding?

     

  3. Data and data quality loops: How can expert human decisions feed back into the system to improve the chatbot and wider service delivery?

 

 

Properly including the humans in the loop 

 

When human agents tackle the tough edge cases and complex enquiries, can you feed what they learn back into your system? Rather than simply close the ticket, use that insight to identify wider organizational problems and, where necessary, invest in fixing system boundary points that affect customer experience.

 

Human judgement is valuable data, either to refine models, adjust guardrails or develop business cases for infrastructure investment. 

 

See human oversight as a continuous source of high quality feedback. Not just customer problems to be fixed, but organizational challenges to be overcome. 

 

When you visualize and design for the whole system, you are properly fixing the system, not just shifting the bottlenecks around. 

 

 

Designing for human judgment 

 

Local quick wins are tempting: they are easy to implement, offer fast metric uplifts and immediate cost reductions. But treating AI as a quick fix patch can move operational friction from one place to another, potentially worsening customer experience or creating costs elsewhere. 

 

Product and experience strategists, data leaders and systems architects need to design resilient systems that consider human expertise as a valuable source of data to keep the entire system grounded. 

 

To build AI capabilities that deliver lasting value, leadership teams need to move beyond automating tasks, deploying quick wins and consider three questions:

 

  1. If we do this, where will the bottleneck shift? Do we understand the wider system well enough to know?

     

  2. What system lag should we expect? How long before we see secondary systemic effects, and are we comfortable with that timeline?

     

  3. How do we best empower the humans in the loop? Can we design work that is varied and sustainable while reducing compliance and customer experience risks and driving wider system improvement?

 

By stepping back and looking at the big picture, organizations can move from fragile patches to human-centric AI systems built for long-term growth.  

Disclaimer: The statements and opinions expressed in this article are those of the author(s) and do not necessarily reflect the positions of Thoughtworks.

Explore a snapshot of today's tech landscape