Enable javascript in your browser for better experience. Need to know to enable it? Go here.

Designing agentic delivery for bounded context and bounded human attention

A coding agent can take a task from its first read of a repository to a working change. In one session it can search the code, draft an implementation, run the tests and open a pull request while a person is still forming the question.

 

Those agents still work inside a finite context window. They still produce a change a person has to understand. As tasks get larger, two limits need designing: the agent’s working context and the reviewer’s attention.

 

 

AI moves the bottleneck

 

Code generation is increasingly less likely to be the slow step.

 

Two constraints become more important:

 

  • the context the model can use effectively.

     

  • the time a reviewer has to decide whether the change matches the ask.

     

Faster generation alone leaves shipping risk unchanged. The rest of this article is about keeping both resources inside bounds.

 

 

Bound the agent’s context

 

Context rot

 

Context rot is a loss of relevance inside a session that still has room left. The useful instruction is still present, but buried under documents, file contents, command output and earlier reasoning the task no longer needs.

 

Example

 

  1. Start. There’s one project rule: never store session tokens in localStorage; use httpOnly cookies. Easy to follow while it is nearly the only instruction in the window.

     

  2. Explore. Mid-session, someone asks how login works today. The agent pulls in a long security guide, greps for localStorage, opens session-manager.ts and a settings page, and pastes a failing test log into the chat. The cookie rule is buried.

     

  3. Drift. Later, the same session adds a “remember me” checkbox. The agent writes the token to localStorage. The original rule is still in the transcript, but it no longer influences the model’s response.

 

A larger context window still leaves the early instruction buried. Capacity and relevance are different problems. What matters is not only whether the instruction fits in the window, but whether it remains salient enough to influence the change.

 

Two habits cause that rot:

 

  • Loading problem: Reading an entire guide to answer one question.

     

  • Exploration problem: Grepping, opening files and pasting test output into the main thread.

     

 

Progressive disclosure

 

The loading problem shows up when the agent pulls an entire guide to answer one question. Most of that text is irrelevant, and it buries the instructions that still matter.

 

Fix the loading problem by giving the agent a map and letting the task decide how far to walk.

 

Example

 

To answer one policy question, the agent used to load all 6,000 lines of business-workflows.md, most of it irrelevant.

 

Instead, split that knowledge into an index and nested flow documents.

 

Question: “What is the SLA exception policy for a stuck approval?”

 

  1. Index. Start at flows-index.md, a short file that points at the rest.

     

  2. Walk. Follow only the path needed:

     

    • workspace-approval flow

       

    • that flow’s escalation rules

       

    • SLA matrix

       

    • exception policy

       

  3. Stop when the answer is found. Five files and a few hundred lines are enough.

     

  4. Leave unrelated material closed. Appeals, asset provisioning and vendor management stay unread.

 

Later questions can still reach any of that material. Up front, load only the index. Discover the rest as needed.

 

 

Sub-agents as context firewalls

 

The exploration problem shows up when greps, file reads and test output land in the main thread. Keeping that exploratory work in the main context can bury the instructions that still matter.Fix the exploration problem by moving the search out of the main session.

 

Example

 

In the cookie-rule session, greps, file reads, and test output used to land in the main thread. That work is real, and costly to keep. It buried the original rule.

 

Instead, hand the exploration to a sub-agent with its own context.

 

  1. Delegate. The sub-agent greps for localStorage, opens session-manager.ts and related files, and reads the failing test log.

     

  2. Isolate. Those matches, logs, and file dumps stay in the sub-agent’s window.

     

  3. Return. A short summary comes back: httpOnly cookies are already implemented in session-manager.ts; reuse setAuthCookie().

     

  4. Keep the conclusion. The main session retains the answer for the next step without carrying the full search trail.

 

What makes the sub-agent useful is its boundary:

 

  • one job

     

  • an isolated window

     

  • a condensed result

     

Exploratory tool calls belong behind that boundary too when the main session only needs their result.Otherwise they sit in the primary context and add no value to the task.

 

 

Bound the reviewer’s attention

 

A person still has to decide whether the change matches the ask, whether the risk is acceptable and whether unrelated edits got bundled in. Agents can produce that change faster than a reviewer can read it. Tests and types can pass while the diff is still too large to judge with care.

 

Evidence of the scale:

 

  • Pull requests regularly run past twenty files and a thousand lines; code volume is up about 30 percent (Salesforce Engineering).

     

  • One industry review found 61 percent of agent-authored pull requests had no recorded human review.

     

  • An example change: 47 files, 3,140 lines, rewriting multiple service repos like auth, billing and onboarding in one go.

 

Cost of a diff that size:

 

  • Bugs sit in the noise.

     

  • Approval becomes a stamp.

     

  • Comments arrive too late to steer the work.

     

  • Unrelated edits are bundled, so a revert is all or nothing.

     

  • Higher change volume has also been associated with higher change-failure rates in some industry data (Cortex, 2026).

     

Line-by-line reading becomes increasingly difficult at that scale. Keep the diffbut add a separate explanation of the work.

 

 

Make changes explain themselves

 

A change that leaves only code behind forces the next reader to interpret why it was made, what it covers and what risk it carries. That reader might be a reviewer, a teammate, a resumed agent or someone debugging months later. With agent-generated code, this gap can widen: volume rises, chat history stays outside the repo and the person who steered the session may not be the one who opens the change later.

 

Carry a short, fixed summary with the work: intent, coverage, how it was checked and what could go wrong. Code stays the source of truth for behavior. The summary holds the decision context.

 

Pull requests are one place to enforce that habit using the same summary structure beside every diff.

 

Example: pull request #482

 

Pull request #482 carries a fixed summary beside the diff. Same shape every time, so a reviewer learns where to look.

 

  • Context. PROJ-482, its blueprint and its readiness note.

     

  • Summary. One paragraph of intent: session handling, invoicing and the onboarding wizard move behind one shared pipeline.

     

  • Changes, each tied to a requirement.

     

    • AC-1: session tokens rotate on password change

       

    • AC-2: invoice totals are recalculated on the server

       

    • AC-3: the onboarding wizard resumes from the last completed step

       

  • Acceptance criterion, test and result. AC-1 → “rotates session token on password change” → passed. Same chain for AC-2 and AC-3: requirement → test → result.

     

  • Risk. Medium: session rotation could log people out mid-rollout. Ships behind a flag; rollback is turning the flag off. No data migration.

     

  • Reviewer checklist. Business logic matches the story. Architecture stays intact. Existing sessions stay as they are until rollout.

 

The summary gives the reviewer a first-pass map of intent, requirement coverage and risk before they inspect the relevant parts of the diff. 

 

Handle mechanical checks and judgment in different places:

 

  • Person confirms decisions with significant blast radius: For example, a plan that edits several repositories at once.

     

  • Automation catches completeness gaps: For example, a missing test for an acceptance criterion and sends the work back before review.

 

The person sees the decision that is costly to undo, and a summary small enough to read.

 

 

Failure modes to design against

 

Long-horizon drift

 

A large story runs across days. Requirements shift in chat, for example, “drop the invoice export; ship token rotation only.” When an agent compacts or summarizes its history, an early plan may survive more reliably than a mid-course correction. The agent can report done against the plan it started with.

 

Store each steering change in the plan file the next session will read. The handoff is a document. A resumed agent picks up the decisions that were actually made.

 

 

Stale specifications

 

Indexes and plan files drift. A document describes a contract the code no longer implements. An agent that trusts the document “fixes” the code toward an abandoned design, for example, reintroducing a deprecated /v1/sessions endpoint because the index still lists it.

 

When code and documentation disagree, the codebase is the authority (for greenfield projects this might differ). A confirmed human decision outranks an agent inference. Revise specs to match the code. Prune ones that no longer match a failure the team has seen.

 

 

Self-verification bias

 

An agent that just wrote the change is a weak witness that the change is done. “AC-1 is covered; tests pass” is a claim.

 

Back that claim with the relevant test run, type check or build result before asking a person to trust it.

 

 

Bounded autonomy

 

The goal is more autonomy, inside limits you set on purpose. An agent can carry a task a long way when these three stay in place:

 

  • Scoped context: Start with an index, open nested docs on demand, keep summaries, discard search trails.

     

  • Short human artifact: Capture intent, requirement-to-test coverage and the risk that needs a decision.

     

  • Mechanical checks first: Run tests, types and builds before involving a person, reserve human attention for judgment with a blast radius.

     

Those limits protect the two scarce resources in agentic delivery: the model’s attention inside the session and the reviewer’s attention after it. Design for them and a larger agentic workflow becomes safer to run.

While the model provides reasoning capacity, harness engineering supplies the structured environment, tool constraints, context boundaries and feedback loops that transforms raw intelligence into a safe, reliable and auditable delivery system.

 


A practical harness has two layers (and selective human gates)

 

To prevent both unconstrained AI drift and human review fatigue, a production-grade harness relies on two complementary enforcement layers, separated by targeted decision gates:

 

 

1. Guides (Pre-action steering): Mechanisms that constrain and direct the agent before it takes action or writes code.

 

 

2. Sensors (Post-action verification): Automated feedback loops that inspect and validate the output after the agent acts.

 

 

3. Selective human gates: Strategic checkpoints placed between guidance and sensors where human judgment or blast-radius sign-off is required.

 

By cleanly separating pre-action steering from post-action verification, teams avoid two common pitfalls: letting agents freelance without constraints or forcing humans to approve every trivial tool call.



Guides: Constraining before the agent acts

 

Guides establish the parameters within which an agent is permitted to explore, plan and write code. Effective guidance relies on four core patterns:

 

Scoped instructions and progressive disclosure

 

Loading monolithic instruction files into every session rapidly fills the context window, causing "attention dilution" and context rot. Instead, instructions should be scoped by file path or domain tree (e.g., loading database rules only when touching schema files). By using progressive disclosure, i.e., traversing documentation trees as needed rather than preloading everything, the agent maintains high attention on relevant team conventions.

 

Least-privilege tools

 

Relying on prompt text to ask an agent not to perform unauthorized actions (such as pushing code or editing configs) is inherently fragile. Least-privilege tools convert behavioral guidance into structural enforcement. For instance, a Q&A or code verification agent is provisioned with read-only tools, literally lacking file-writing capabilities.

 

Explicit defaults instead of guesses

 

When an agent encounters optional inputs during structured intake (such as an optional technical spec), it should avoid making unstated assumptions. The harness enforces an "ask before deciding" protocol. If the developer chooses to skip an optional step, the harness can explicitly record that decision. A default is visible and inspectable, a guess is hidden and risky.

 

Confirmation only for high-blast-radius decisions

 

To maintain high velocity without sacrificing safety, human confirmation gates are reserved for irreversible or high-blast-radius choices, such as multi-repository scope changes, schema migrations or public API modifications. Lower-risk decisions proceed with sensible defaults.



Sensors: Verifying after the agent acts

 

Even the best guides cannot catch every logic edge case or runtime error. Sensors operate after the agent acts, providing automated computational feedback to validate output.

 

Automated verification

 

Sensors wrap standard software engineering checks like unit tests, linters, static type checkers and architecture rules into the agent loop. When an agent finishes an implementation pass, sensors run automatically to evaluate correctness rather than relying on the agent's self-assessment.

 

Silent success, verbose failure

 

Sensors follow the "silent success, verbose failure" rule:

 

  • When checks pass: The sensor produces minimal output, preserving valuable context tokens and human attention.

 

  • When checks fail: The sensor generates verbose, actionable feedback (stack traces, linter error locations, failing test diffs), feeding the error directly back into the agent loop so it can self-correct without human intervention.

 

Promoting rules from prose into executable enforcement

 

If a team finds themselves repeatedly adding prose rules to prompt files because an agent ignores a convention, that rule should be promoted into a mechanical sensor (a custom linter rule, a type check or an architecture test). Prose is the starting point; where a rule can be reliably encoded, mechanical enforcement is the stronger option.



How the pattern works across repositories

 

To illustrate how a shared harness can help, consider a system with three microservices:

 

  • billing-service (owns a shared discount_rate field)

 

  • checkout-service (consumes discount_rate)

 

  • invoicing-service (consumes discount_rate)

 

This simplified example illustrates how the pattern can work when dependencies are known and accessible to the harness.

 

Unconstrained repository-level execution

 

When an agent is executed within billing-service without a multi-repo harness, a refactoring request causes the agent to rename discount_rate to promotional_discount in billing-service alone. The agent reports success because its local tests pass. However, checkout-service and invoicing-service silently drift out of sync, breaking production workflows.

 

Shared harness execution

 

When the same request is processed through a shared harness wrapping all three repositories:

 

  1. Impact analysis: The agent scans dependency graphs across the shared workspace and identifies that changing billing-service affects checkout-service and invoicing-service.

     

  2. Multi-repo confirmation gate: Recognizing a multi-repo blast radius, the harness pauses execution and presents a blocking confirmation gate to the developer.

Disclaimer: The statements and opinions expressed in this article are those of the author(s) and do not necessarily reflect the positions of Thoughtworks.

Explore a snapshot of today's tech landscape