The Agentic Harness: How We Rebuilt Our SDLC for AI-Augmented Product Engineering
Prakash Donga•8 Oct 26•7 Min Read

A few months ago, a small Tailwind config change broke position: sticky across our nav and header layer. The root cause was three lines of code.
The deeper problem was that the fix lived in one engineer’s head instead of somewhere every future engineer and AI agent would read before changing the codebase.
That incident changed what we optimized for. We needed production lessons, implementation constraints, and project context to live somewhere both engineers and agents could use before touching the code.
We call that operating system the Agentic Harness: a structure for context, specifications, execution, and verification around AI-generated work.
This is how we rebuilt our SDLC around it.
What Is AI-Augmented Product Engineering?
AI-augmented product engineering is a development model where AI participates directly in planning, building, and verifying software within defined engineering constraints.
It sits between two related models.
- AI-assisted development uses tools such as autocomplete and chat to speed up individual engineering tasks. The engineer still carries most of the project context and decision-making.
- AI-augmented development gives AI a larger role across requirements, planning, implementation, testing, and review, with defined human checkpoints around important decisions.
- AI-native development is the direction this model moves toward: agents become first-class participants in delivery, while human involvement concentrates around judgment, risk, and accountability.
At SoluteLabs, we currently operate with AI throughout the lifecycle, but we remove human gates only when the evidence from repeated use says it is safe to do so.
The Agentic Harness is what makes that process repeatable.
Why We Rebuilt Our SDLC
We rebuilt the process because four problems kept showing up.
1. Slow development cycles
Engineers repeatedly had to rediscover old decisions, workarounds, client constraints, and previous fixes because that context was not captured in a durable system.
2. Manual testing bottlenecks
Visual regressions across a large component system are difficult to catch by manually clicking through pages before every release. Too often, the first signal of a regression came after deployment.
3. Fragmented AI workflows
Engineers were using different combinations of Copilot, ChatGPT, manual search, and ad hoc review processes.
There was no shared standard for planning, reviewing, or recording what an AI-assisted change had actually verified. That made PR review slower because reviewers had to reconstruct what was checked and what was assumed.
4. Senior engineering time was spent on co-ordination
Experienced engineers were reconstructing context, reviewing first-draft output, and rechecking routine changes instead of spending that time on architecture, risk, and product decisions.
What the Harness Is Designed to Prevent
The Agentic Harness was shaped around the failure modes we kept seeing as AI-generated code moved closer to production.
1. Context hallucination
The agent does not know why an unusual implementation exists, so it “cleans up” code that was protecting the system from a known edge case.
2. Self-confirming code
The same agent writes the code and validates it against its own assumptions. The implementation looks coherent, but the validation only confirms the original misunderstanding.
3. Silent regressions
A change works locally but breaks an unrelated UI state, layout rule, webhook path, or environment-specific behavior.
4. Verification debt
AI can increase implementation speed faster than a team can validate the output.
We think of that gap as verification debt: the distance between how quickly AI can produce code and how confidently the team can prove that code is safe to ship.
That became the main problem the harness needed to solve.
The Agentic Harness in Practice
The harness runs through four phases:
INFORM → SPEC → BUILD → VERIFY
Each phase produces a defined artifact and has a clear human checkpoint.

01. INFORM: Context
Before implementation starts, the harness assembles the context needed to understand the work.
- Requirement synthesis
Raw inputs such as client briefs, bug reports, and Slack discussions are converted into structured requirements with unique IDs and a traceability matrix.
That gives later implementation work a direct link back to a named requirement rather than a vague ticket title.
- Pattern recognition from past project data
At the start of a new phase, we carry forward the open items, known constraints, deferred work, and implementation lessons from the previous milestone.
The agent starts with inherited project context instead of reconstructing it from scratch.
- CLAUDE.md as operational memory
CLAUDE.md is checked into the repository and read by both engineers and AI agents at the start of a session.
It records implementation constraints that are easy to miss from the code alone.
Examples include:
- why a specific @reference directive can trigger an out-of-memory build failure
- the fallback chain required for a Sanity dataset environment variable
- a useEffect cleanup pattern that can break position: sticky across the site
Many of those entries exist because we learned the constraint through a real failure.
Recording them in CLAUDE.md makes the same failure much harder to repeat.
Human checkpoint: validate scope and requirements.
02. SPEC: Planning
Once the context is assembled, the harness produces a plan before implementation begins.
- AI-assisted system design
For a new subsystem, AI can produce multiple implementation approaches with explicit trade-offs rather than a single recommendation.
- Trade-off simulation
Before committing to an approach, we ask the model to challenge it.
- What breaks under load?
- What happens after a schema change?
- What fails if a third-party API changes its contract?
This exposes assumptions before they are buried in code.
- Spec-driven development
Each phase gets written planning artifacts before implementation starts:
- project scope
- requirements with REQ-IDs
- phased roadmap
- architecture decisions where required
The plan can be reviewed and rejected before implementation begins.
Architecture, irreversible trade-offs, and build-versus-buy decisions remain with the engineering lead.
Human checkpoint: approve architecture and trade-offs.
03. BUILD: Execution
This is where the coding agent executes against the approved specification.
In our current stack, that agentic coding tool is Claude Code.
- Code generation
The agent implements against the existing codebase and project conventions rather than producing isolated boilerplate.
- Refactoring legacy systems
The same workflow can be used for migration work.
For example, a repetitive configuration pattern spread across many files can be consolidated into a single source of truth, with the agent handling the mechanical changes while review focuses on the higher-risk parts.
- Independent agent review
For non-trivial changes, we do not rely on a single model pass.
One agent executes against the approved specification. A separate review pass checks the diff for missed constraints, ordering assumptions, guard placement, environment-variable handling, and other class-of-bug issues before human review.
This is the same multi-agent orchestration we build for clients, running on our own codebase first. In our multi-agent knowledge platform for a Fortune 500 retailer, the same instinct shows up in the product itself: answers below a 0.38 confidence threshold are refused, and plans are validated against schemas before they execute.
- AI-assisted deploy verification
Deployments remain human-triggered through a defined checklist.
The proof step is automated through production probes, status-code checks, and response-body assertions that confirm the intended change reached production and detect obvious regressions.
Human checkpoint: review the diff and risk areas.
04. VERIFY: Verification
A change does not ship because the implementation agent reports that it works.
- Automated test generation
Visual regression tests are generated and run against a baseline before and after changes, including the class of layout issue that originally broke the sticky navigation.
- AI-assisted PR review
A dedicated review pass, separate from the implementation agent, checks the diff for correctness issues, unnecessary complexity, and reuse opportunities before a human reviewer sees it.
The human starts from a reviewed diff rather than the first draft.
- Security and compliance checks
Changes touching data queries or external webhooks go through a dedicated security review pass.
For example, a query layer can be hardened by replacing string interpolation with parameterized queries, then verifying that no unparameterized call sites remain.
- Production regression checks
We maintain a short set of production probes and revert-detection canaries to catch cases where a later deployment silently undoes an earlier fix, especially when multiple branches pass through the same environment.
Human checkpoint: sign off on evidence before release.
Before vs. After: What Actually Changed
| Dimension | Before | After |
|---|---|---|
Cycle time | Context rediscovered on new work | Project context loaded from CLAUDE.md and prior-phase handoff docs |
Defect detection | Client reports and manual QA | Visual regression and production probes |
Deployment pattern | Larger, batched releases | Smaller changes with verification |
PR review | Human reviews first draft | AI review first, human reviews a cleaner diff |
Senior engineering time | Reconstructing context and checking routine changes | More time on architecture, risk, and client-specific work |
The important change is not one individual tool.
Context, execution, and verification now follow a repeatable system.
What We Learned
Write constraints down as soon as they become known.
Undocumented constraints are where agents make plausible but incorrect changes.
If a production issue reveals something the code alone cannot explain, that lesson needs to become part of the project context.
Standardize the workflow before standardizing every tool.
The exact model or coding assistant may change.
The more important standard is how work is scoped, reviewed, verified, and traced back to requirements.
Keep human judgement around irreversible decisions.
Architecture, security-sensitive changes, major trade-offs, and release approval still need named ownership.
AI can accelerate execution without taking ownership of accountability.
What This Changes for Clients
For clients, the harness makes delivery easier to inspect.
Requirements are broken into unique IDs and mapped through a traceability matrix before a phase starts.
AI-generated changes move through review evidence such as independent review, human sign-off, and deployment probes before release.
Visual regression and production checks run in the delivery pipeline before the change reaches a client-facing environment.
That gives senior engineers more room to focus on architecture, edge cases, and product judgment.
Conclusion
The bigger change is that context and verification no longer have to live inside individual engineers’ heads.
They can become part of the delivery system itself.
That is how we reduce verification debt as AI-generated code moves closer to production: by making the work traceable, reviewable, and easier to validate before it ships.
If your team is moving from AI-assisted coding toward a more structured AI-native delivery model, we can help design that system around your product, codebase, and release process.
