● Engineering

Agents Create Leverage. Harnesses Create Trust.

Over the last year at Travelopia, we have seen our AI-enabled delivery practice slowly mature. For me, maturity does not mean a few impressive demos or a handful of enthusiastic early adopters. It means the engineering team adopts a version of the process, uses it consistently, and delivers with it over months. That is the real test.

We have been iterating on AI-enabled delivery for a while now, and it has matured. It works. It has helped teams build more confidence with this way of working. But naturally, the next question has started appearing: what does the next step look like?

That question led me, in my personal experimentation, toward autonomous agents. It is still experimental. I am still learning. But tinkering with agents made one thing very clear: the agent itself is only one part of the story. The bigger question is the system around it.

Everyone asking about AI agents starts with the same question: “Can the agent do the task?” For an engineering leader, that is the wrong first question. The one that matters is: “Can we safely trust it with this responsibility?”

That difference matters because the jump from AI-enabled delivery to agent-owned delivery is not just a tooling upgrade. It is an operating model upgrade. With AI-enabled delivery, humans are still driving the flow, but AI is embedded into how the team plans, codes, tests, reviews, and delivers work. With agent-owned delivery, we start asking software to take ownership of bounded tasks: understand the requirement, choose steps, use tools, modify code, verify work, and prepare something for review. That creates leverage. But leverage without trust does not scale. That is where the harness comes in.


The harness is the control layer around the agent

An agent gives us leverage. It can move faster than a human in many parts of the software delivery process — reading a codebase, making changes, writing tests, fixing errors, preparing a pull request. That is powerful. But the harness decides whether that power can be trusted.

In this context, a harness is not a single tool. It is the control layer around the agent — everything that makes it usable in a real team:

  • What context it receives
  • What tools it can access
  • Which repositories it can read
  • Which files it can modify
  • Whether it works inside a sandbox
  • How it runs tests
  • When it must ask for clarification
  • When a human must approve
  • What logs are captured
  • How failures are rolled back
  • Who is accountable for the final decision

Without this, an agent is just raw capability. Useful? Yes. Safe at scale? Not necessarily.

Diagram with an agent at the centre surrounded by control rings representing context, tools, permissions, sandboxing, tests, human approval checkpoints, logs, and rollback paths — together forming the harness
The agent creates leverage, but the surrounding controls decide whether that leverage can be trusted.

Better models increase the ceiling. Better harnesses raise the floor.

It is tempting to think the main advantage will come from choosing the best model or the most advanced agent framework. And yes, models matter — a better model can reason better, understand ambiguity better, generate better code, and recover from mistakes faster.

But in an engineering team, the bigger operational question is not only what is possible. It is what is reliable. That is where the harness becomes critical.

A weak model inside a strong harness may be limited, but at least it is contained. A strong model inside a weak harness may be impressive, but it can also become dangerous. The model determines the ceiling. The harness raises the floor. And for leaders, the floor matters a lot — it is what decides whether the system is safe enough for repeated use by real teams on real codebases.


“Human in the loop” is not a complete strategy

A common answer to agent risk is simple: “Keep a human in the loop.” That sounds reasonable, but it is not enough. Where, exactly, is the human in the loop? Before the task starts? During planning? Before code changes? Before a pull request? Before deployment? After something breaks?

Human oversight needs design. Otherwise it becomes a vague safety phrase. In traditional software delivery, we do not rely on trust alone — we rely on branches, reviews, tests, CI checks, environments, access control, logs, alerts, and rollback paths. Agents should be no different. In fact, they need this discipline even more, because agents make execution faster. And when execution gets faster, weak controls become visible very quickly.


Autonomy is not a switch. It is a ladder.

One mistake teams make is treating agent adoption as a binary decision: either we use agents or we do not, either we trust them or we do not. That is not how engineering organizations should think about this.

Autonomy should be a ladder. At the first level, an agent may only suggest an approach. Then it may draft code. Then it may make local changes, run tests, open a pull request, and eventually fix bounded issues end to end. For some narrow and well-understood areas, it may operate with more independence. But each step up the ladder requires a stronger harness.

More autonomy should not come from excitement. It should come from evidence:

  • Can the agent stay within scope?
  • Can it explain what it changed?
  • Can it run the right checks?
  • Can it identify uncertainty?
  • Can it stop when the task becomes risky?
  • Can it produce a clean, reviewable pull request?
  • Can the team trace what happened?

If the answer is no, the agent is not ready for the next level of autonomy.

Ladder diagram with rungs labelled from bottom to top: suggest approach, draft code, make local changes, run tests, open pull request, fix bounded issues end to end, operate independently — each rung annotated with stronger harness requirements
Each level of autonomy requires stronger evidence, stronger controls, and clearer accountability.

The real leadership work is deciding what the agent should not do

Most AI conversations focus on capability: can it write code, fix bugs, understand architecture, talk to APIs, deploy? But mature leadership also asks the opposite set of questions. What should the agent never touch? Authentication flows? Payment logic? Infrastructure? Secrets? Database migrations? Architecture decisions? Production deployments?

These boundaries are not signs of mistrust. They are signs of responsible delegation. Good leaders do not give humans unlimited access on day one either — we define roles, permissions, review processes, escalation paths, and accountability. Agents need the same thinking. Maybe even stricter.


Agents reward boring engineering discipline

There is an interesting side effect: agents make boring engineering practices more valuable. A clean README matters more. Good tests matter more. Consistent folder structure, clear architecture decisions, useful documentation, small pull requests, well-written tickets — all of it matters more.

Because agents depend heavily on the environment we give them. A messy codebase with poor tests and unclear conventions makes an agent unpredictable. A disciplined codebase gives the agent a better chance of producing useful work. AI does not remove the need for engineering hygiene. It increases the return on it. The teams that already invested in clarity, automation, and discipline will get more out of agents than teams hoping agents will magically fix their chaos.


Buying the tool is not the strategy

Buying a tool is not the strategy. Adding a coding agent to the SDLC is not the strategy. Letting developers experiment with AI is not the strategy. The strategy is designing how AI-enabled delivery should work inside the organization — what tasks are safe for agents, what input quality is required, who writes the task brief, what context and permissions the agent gets, what checks must pass, who reviews the output, what happens when it is wrong, and how the system improves over time. This is not just an engineering tooling conversation. It is an operating model conversation.


Is the harness just temporary scaffolding?

A fair counterargument is that today’s harnesses may look necessary only because agents are still immature.

As models get better, they will need less hand-holding. They will understand context better, recover from mistakes faster, ask better clarification questions, and make fewer careless changes. Some checks that feel necessary today may become lighter over time.

I think that is true.

But that does not make the harness temporary.

The harness is not only there to compensate for weak models. It is there to define accountability around delegated action.

Even if agents become dramatically more capable, an organization still needs to decide what they can access, what they can change, what evidence they must produce, when they must stop, and who owns the final decision.

Better models may reduce friction. They do not remove responsibility.

The control layer will evolve, but it will not disappear. Trust is not just a model capability. Trust is a system design choice.


The future belongs to teams that can delegate safely

I do believe agents will become a serious part of software delivery — not because they are perfect, but because even imperfect agents can create meaningful leverage when the work is bounded correctly.

The winning teams will not be the ones that give agents the most freedom. They will be the ones that create the clearest trust boundaries — knowing where agents can move fast, where humans must decide, where automation should verify, and where the system must stop. That is the real work ahead for engineering leaders. Not just adopting agents. Not just comparing models. Not just running impressive demos. But designing the harness that makes autonomy useful, safe, and scalable.

Agents create leverage. Harnesses create trust. And without trust, leverage does not scale.

· · ·
Niraj Chauhan
Niraj Chauhan writes here about software craftsmanship, running, and family life in Bengaluru.