Software delivery became workflow-driven for a reason. As organizations grew, work spread across developers, product managers, security teams, operations teams, and increasingly specialized tools. We created issue states, pull requests, CI pipelines, approval gates, deployment stages, and handoffs to coordinate that work. Then we surrounded those workflows with ceremonies to keep everyone synchronized: sprint planning, standups, reviews, demos, retrospectives.

Those ceremonies are not arbitrary overhead. They are synchronization mechanisms for a human-driven system where the state of a software change is fragmented across people and applications. Jira knows the issue. GitHub knows the code. Jenkins knows the build. Humans carry much of the context connecting those records together.

It is natural that DevOps vendors are now adding agents to the workflows they already have. Put an agent into a step. Let it implement an issue, investigate a failure, review a pull request, or recommend what happens next. The existing workflow remains the control plane.

We are automating the participants without questioning the ceremony.

Intent as Code does not prescribe the path

Agents introduce a different requirement. We can know what we want to accomplish without knowing exactly how the work will get there.

This is where Intent as Code becomes more interesting than storing another specification next to the source code. Intent should be a durable, addressable primitive that captures why a change needs to exist, what conditions define success, and enough structure for a machine to determine whether the resulting work actually satisfied it.

We saw this resonate recently at WeAreDevelopers when we demonstrated how Atomic starts with intent. The why and acceptance criteria are established first. Tasks, changes, evidence, and the implementation itself can emerge underneath them as the agent learns what the work requires.

The intent is known. The workflow is not.

An agent may inspect one subsystem and discover a dependency in another. It may test an assumption and learn that it was wrong. It may delegate a bounded intent to a child agent, which can discover more work and delegate again. Each child carries its own intent, gathers evidence, and eventually returns its result to the parent.

The structure repeats:

intent → reasoning → child intent → work → evidence → reconciliation

What remains stable is not the sequence of steps. It is the intent against which the resulting work must eventually reconcile.

Trying to encode every possible reasoning transition into a workflow defeats part of the reason we use an agent. But allowing an agent to discover the path does not mean that path needs to disappear when the session ends.

Agent harnesses already expose lifecycle events through tool calls, hooks, plugins, effects, sessions, and sub-agents. If we capture the meaningful events and preserve their causal relationships, the architecture begins to invert.

Intent as Code makes intent the instruction. The workflow becomes the evidence.

The workflow is discovered during execution, but it can still become durable afterward.

The graph learns while the agent works

A durable causal history tells us more than what steps executed. It tells us why they happened, what changed as a result, and what evidence connects the result back to the original intent.

It can also teach the system the language of the software it is changing.

We have spent decades building search systems that work well once we provide enough structure. Technologies such as Lucene made enormous bodies of information searchable, while fields, analyzers, synonyms, taxonomies, and domain dictionaries helped us teach those systems how to interpret the vocabulary they indexed.

The difficult part was meaning.

An ecommerce application speaks in products, carts, orders, inventory reservations, payments, fulfillment, and refunds. A healthcare application speaks in patients, encounters, observations, medications, orders, and providers. Trying to define the complete vocabulary before agents begin useful work recreates the same mistake as trying to define their complete workflow beforehand.

Agents let us build the dictionary while learning the language.

An agent may discover InventoryReservation because an intent causes it to trace the relationship between checkout and inventory. Another intent may later connect that same concept to fulfillment. Those relationships become durable because they were discovered while performing actual work, not because somebody anticipated every domain concept in advance.

The ontology can provide the grammar. The causal history builds the dictionary.

That distinction also changes how context can be assembled. Expensive reasoning can discover a genuinely new concept or relationship once. A classifier can then map subsequent changes against vocabulary the system already understands. Those classifications can drive semantic views, giving the next agent the causal closure relevant to its current intent instead of handing it an entire repository and asking it to rediscover the domain.

The work of one agent becomes context for the next.

Common meaning matters more than common reasoning

We have solved a version of this interoperability problem before.

When my team worked on Probot at GitHub, Jenkins, Jira, GitHub Apps, and other systems did not need to share an implementation. They needed common event boundaries that allowed independently built systems to participate in GitHub's lifecycle.

Agent harnesses present a similar opportunity.

Claude Code can retain its lifecycle. Codex can have another. OpenCode can expose its own hooks, plugins, and effects. Whatever comes next will make different choices again.

We should not standardize how those systems reason.

We need to normalize what their events mean in the lifecycle of software change.

One event may mean an intent was created. Another may represent evidence gathered against an acceptance criterion. Another may identify a change. Another may tell us that a child intent believes its work is complete.

The harness owns how the agent works. The common layer establishes what happened, why it matters, and how it relates to the state we already know.

Agent interoperability does not require standardized reasoning. It requires standardized meaning at the boundaries of change.

The graph has to run in both directions

The forward traversal is only half of the architecture.

As an agent works, the causal graph expands outward. Intent establishes why the work exists. Acceptance criteria define what must become true. Tasks and child intents emerge as the problem is decomposed. Changes and evidence record what actually happened.

Triage makes the inverse traversal.

Start with the change. What evidence supports it? Which task authorized it? Which acceptance criterion did it satisfy? Which intent did that criterion belong to? Does the resulting state actually satisfy the reason the work began?

Execution expands the graph outward. Triage walks the graph back toward intent.

If the graph reconciles, the intent can close. If it does not, the unresolved state becomes the beginning of another iteration. That gives us a more useful definition of completion: an agent is not done because it stopped producing tokens or produced a patch. It is done when the resulting state reconciles with the intent that caused the work.

That freedom still requires deterministic boundaries. Cost, tokens, elapsed time, recursion, tool use, and permissions need hard limits enforced outside the model. Nondeterministic reasoning cannot mean unlimited consumption.

There is another boundary when reasoning turns into an external effect. Deploying an artifact, provisioning infrastructure, requesting an approval, or performing a transaction introduces requirements around retries, idempotency, authorization, persistence, and failure semantics.

This is where workflow begins.

One deterministic boundary controls what the agent can consume. The other controls what it can affect. Between those boundaries, the agent is free to discover the path.

The workflow the agent leaves behind

This is the architecture we are building toward with Atomic.

Atomic treats Intent as Code as part of the causal definition of software change. Agent harnesses keep their own reasoning and effect mechanisms, while Atomic connects intent, acceptance criteria, tasks, domain concepts, changes, evidence, provenance, attestations, and triage into a durable graph.

Agents walk forward from intent. Triage walks backward from change. The graph learns along the way, and semantic views can give the next agent better context than the one before it had.

Human-driven DevOps needed us to define workflows first and continually synchronize people around them. An agent-native system gives us another option: start with durable intent, let agents discover the work, preserve their lifecycle events, and use the resulting causal history to understand both what happened and whether it accomplished what we asked for.

The workflow is not the graph we draw before the agent starts. It is the causal graph the agent leaves behind.

Common questions

What does intent as code mean for AI agent development?

Intent as code means an agent starts from why a change needs to exist and the acceptance criteria that define what must become true. The agent discovers the path, including any child intents it delegates, and Atomic records the resulting work and evidence. Intent becomes the instruction. Workflow becomes the evidence.

Why not just add AI agents to existing DevOps workflows?

Existing workflows and ceremonies exist to synchronize humans across fragmented tools like Jira, GitHub, and Jenkins. Dropping an agent into a step automates the participant without questioning the ceremony. Agents often discover dependencies or invalidate assumptions mid-task, so the stable control point should be the intent and its acceptance criteria, not a predetermined sequence of steps.

How can I trace what an AI coding agent did across tools like Claude Code, Codex, and OpenCode?

Capture the lifecycle events each harness already exposes, such as tool calls, hooks, sessions, and sub-agents, and normalize what those events mean for a software change. Atomic does not standardize how agents reason. It records the causal relationships between intent, tasks, changes, and evidence, so the path an agent discovered stays durable after the session ends.

How do I give AI coding agents memory so they stop repeating mistakes and burning tokens?

Build context from the causal history of past work. As agents traverse the codebase, Atomic learns domain concepts and their relationships. Expensive reasoning discovers new concepts, and cheaper classification maps later work against that vocabulary. Each new intent gets only the context related to it, instead of the whole repository.

How do I know when an AI coding agent is actually done?

An agent is done when the resulting state reconciles with the intent that caused the work, not when it stops producing tokens or produces a patch. Triage walks backward from the change to its evidence, authorizing task, acceptance criterion, and intent. If the graph does not reconcile, the unresolved state starts another iteration.

How should engineering teams govern AI coding agents?

Set two deterministic boundaries. One controls what an agent can consume: cost, tokens, time, recursion, tools, and permissions, enforced outside the model. The other controls what it can affect: deployments and other external effects need authorization, retries, idempotency, persistence, and failure semantics. Between those boundaries, the agent is free to discover the path.

Build on a foundation that remembers.

Install the CLI and start recording from the next agent turn. No account required.

Install Atomic → Read the docs