AI adoption is done, and yet the results stay in individual hands. What separates one organization from another is not model performance. It is whether they have a container that gives AI its purpose, its constraints, what to refer to and what gets checked — a Harness.
One container, though, is not enough. Upstream, in experience strategy and touchpoint planning, and downstream, in content, design and code, each needs its own. Deciding how the vertical and the horizontal connect is what an Agent Harness Architecture does.
This article runs from the definitions we actually use, through the structural source of truth that connects upstream to downstream, to where we place the gates that people decide at and the ones machines run.
Contents
- You adopted AI, but nothing accumulates
- Harness and Agent Harness Architecture
- The eight things a Harness must have
- A Design Agent Harness alone cannot build a touchpoint
- The Standards / Knowledge Layer every Harness refers to
- What connects upstream and downstream is a structural source of truth, not prose
- Separate Standards, Findings and Gaps
- Place Human Gates and Quality Gates in advance
- Writing the definition and keeping the working parts are two different things
- The moat is in Experience Intelligence
You adopted AI, but nothing accumulates
Most companies have finished the adoption part. Accounts for everyone, a shared prompt library, internal workshops. Individual work did get faster. And yet, when someone asks whether this has become an organizational result, the answer gets vague. We hear this often from people responsible for DX and AI.
Why does it happen? Because the people who are good at using AI are simply good at it, and their way of working has not become part of how the company works. A skilled person pours the purpose and the constraints in their head into a prompt each time, and judges the output with their own eyes. Craft work has picked up a fast tool and become faster craft work. That is all.
What separates organizations here is not model performance. With the same model, some organizations accumulate results and some do not. The difference is whether AI is given a container that tells it what the goal is, what constraints apply, what to refer to, and how far it may decide on its own — and whether those containers are designed to work together. The first is called a Harness, and the term is already common in the AI agent field. Our point is what comes after it: placing Harnesses both vertically and horizontally and designing how they connect. That is what we call an Agent Harness Architecture.
Harness and Agent Harness Architecture
A harness is originally a piece of tack for a horse. It does not make the horse stronger. It transmits the horse’s power in the direction you want to go, without breaking anything. Software has its own older sense of the word — a test harness runs code and checks the result against what was expected — and the AI usage descends from both. Either way, it is a device that directs and verifies power rather than adding to it.
Some definitions. An Agent has a limited role and produces a judgment or an output from a given input. Agent here does not mean only the autonomous kind that keeps deciding and acting on its own; it also covers one that performs a bounded step inside a defined sequence. A Skill is a reusable unit of work an Agent calls — a competitor scan, a terminology check. Note that Skill is also the name of a specific packaging format in current use; what we mean here is the unit of work, not that format.
A Harness gives an Agent its purpose, context, constraints, reference knowledge, tools, verification, Human Gates, and feedback paths, so that it works toward the goal of the project. The common shorthand is Agent = Model + Harness, where the Harness covers everything outside the model’s reasoning: tool execution, memory, state, the execution environment, and safety controls. Our definition adds the project-design elements to that. Orchestration is the layer that launches multiple Harnesses and manages dependencies, sequence, branching, integration, and rework.
Here is the point. Building one Harness and having an Agent Harness Architecture are not the same thing.
A Harness is a container. It lets AI produce output that is consistently yours within one domain. That is useful, but on its own it stays a point. An Agent Harness Architecture selects the Harnesses a project needs, decides dependencies and sequence, places the Standards they all refer to, checks output at Quality Gates, decides which layer a failure goes back to, and returns what was learned to the next project. It is the design of the whole.
To keep the metaphor: the Harness is the tack, the Architecture is the carriage, the driver, and the route. However many good pieces of tack you own, the load does not arrive on its own. When adding more tools does not make results accumulate, the missing piece is usually not another container. It is the design that connects them.
The eight things a Harness must have
A Harness is not a prompt library. A prompt is an instruction. A Harness also includes the mechanism that asks back when the input is insufficient, the mechanism that checks the output by machine, and the path that says who a failure returns to. We define eight things a Harness must have at minimum: Input, Output, Questions, Process, Standards, Quality Gate, Risk / Guardrail, and Learning.
The two that go missing most often in practice are Questions and Learning. Without the first, generation runs on incomplete input. Without the second, neither success nor failure reaches the next project. And the very existence of somewhere to return learning to presupposes an Architecture.
A Design Agent Harness alone cannot build a touchpoint
The Harness that gets the most attention right now is the one for design: putting your design rules and criteria into a form AI can refer to, so that it reliably produces work that looks like yours. That is a worthwhile thing to do.
But it does not, by itself, build a customer touchpoint. What a Design Agent Harness guarantees is consistency of what comes out. If no one has decided whose state should change, in which direction, and which touchpoint carries that change, then what you get is work that is consistent and off-target, produced quickly and in volume.
Worse, when the upstream is vague and you let the downstream run, that vagueness does not disappear. It comes back amplified as variance. If the Experience Strategy admits two readings, the touchpoint design splits into four, and the deliverables into eight. When people did this work, an experienced practitioner absorbed the variance along the way. AI does not absorb it. Within the scope it was given, it produces several plausible things.
So Harnesses are needed vertically, upstream as well. An Experience Strategy Harness reads market, competition, customers, brand, business goals and constraints, and decides whose state should change and in which direction. A Touchpoint / UX Planning Harness converts that conclusion into journeys, the role of each touchpoint, information architecture, functional and content requirements, and evaluation criteria. What an upstream Harness produces is not a deliverable. It is the criteria the downstream will judge by.
And vertical is not the only gap. Horizontally, inside the same Production Harnesses layer, a Design Agent Harness needs a Content Agent Harness and a Coding Agent Harness beside it. The Content Agent Harness refers to Brand Voice, evidence, rights, channel constraints and SEO / GEO requirements, and verifies consistency and factual accuracy. The Coding Agent Harness refers to Engineering Standards, architecture, security, performance, accessibility, testing and deployment conditions, and iterates generation, review, testing and correction. The Standards they refer to and the way they are checked are different, so each needs its own container.
These three are also not a serial handoff. Content is constrained by layout, Design by implementation feasibility, and the results of Coding feed back into Design and UX. Precisely because the relationship is bidirectional, all three need one shared object to read.
Harnesses, then, are needed both vertically and horizontally. Deciding how the vertical and the horizontal connect is what an Agent Harness Architecture does.
The Standards / Knowledge Layer every Harness refers to
What upstream Harnesses and the extraction process produce has names: Design System, Content System, Engineering Standards. The Design System covers tokens, a component inventory, layout rules and breakpoints. The Content System covers glossary, canonical notation, prohibited terms, voice, slot length ranges, and metadata and URL conventions. Engineering Standards cover stack, directory structure, naming, implementation conventions, and performance and accessibility criteria.
Together these three make up what we call the Standards / Knowledge Layer. It is not a Harness. It is the source of truth that every Harness refers to. Downstream Agents decide by reading this layer rather than by improvising.
The layer holds two kinds of content, kept apart. Client- and project-specific material — business, brand and user context, UX strategy, journeys, contractual, rights, security and AI-use conditions, and the project’s own decision log. And material common to Neuromagic — the shapes of Modules, Skills and Harnesses, evaluation criteria, review procedures, the format for prohibited patterns, reusable templates and guardrails, and patterns for classifying projects and judging estimates, staffing and risk. Keeping the two from mixing, and enforcing that separation as a mechanism rather than as a stated intention, is what pays off later.
What connects upstream and downstream is a structural source of truth, not prose
Consider building a website. If the only thing handed to the Content, Design and Coding Harnesses is a brief that says “please make a page like this”, what happens? Prose cannot fix the number of slots, their order, or their length, so the three imagine three different structures. People then compare the outputs and reconcile them. The faster you produce, the faster you diverge.
So we place a structural source of truth in the Touchpoint / UX Planning Harness layer, for all three to read. The format is HTML with no styling applied at all. It carries elements, hierarchy and order, and nothing about color, typeface or spacing. It does not fix the wording either. What it fixes is that this heading has this many elements under it, each within a stated character range.
Why not fix the wording? Because the moment you do, the upstream has pre-empted a downstream decision and the boundary of responsibility disappears. Then why fix the character range upstream? Because the range determines layout, so it affects both Design and Coding — which makes it part of the structure. It looks like a small distinction. In practice it is the line that stops the team from relitigating “where does structure end and visual design begin” on every project.
This is not specific to the web. In any business process, if you do not fix a shared object before handing work to AI, the smarter the model, the more confidently it will build something impressive on its own reading.
Separate Standards, Findings and Gaps
A word about existing assets. Renewals, operational improvement and incremental development are more common than greenfield work, and there the starting point is an existing site, app or brand guideline. That is the job of the Standards Extraction Harness: to derive the Design System, Content System and Engineering Standards from existing touchpoints.
The line this Harness holds is the separation between what is canonical and what merely exists. Not everything that is running is correct. Inconsistent notation, accessibility violations, values that were never tokenized. Take them in without sorting them, and you reproduce the client’s existing defects in your own deliverables. So the output is split into Standards, Findings (defects in the current asset) and Gaps (areas that were not present in what was observed). The third matters most: do not fill an unobserved area by inference, state explicitly that it needs new design.
Every value also carries a Provenance Ledger entry: its origin (declared, measured or inferred), its scope (project or shared), source URL, date and conditions of observation. More than the value itself, what decides later judgment is whether it was declared by the client, rounded from our own measurement, or inferred. And without scope recorded, once a project ends nobody can say whether a value may travel to the next one.
Place Human Gates and Quality Gates in advance
“How much should we hand over to AI” always comes up. Our answer is not a percentage. It is to design in advance where people decide.
The dividing rule is clear. Anything that cannot be settled without business context does not go to a machine. Whether a defect in a current asset is something to fix or a deliberate client choice cannot be decided without knowing the client’s business priorities. That is a Human Gate. Whether required slots are filled, whether prohibited patterns crept in — that is a Quality Gate, run by machine. You run it first so that people are not asked to spend time on work that has already failed.
A Human Gate records three things: the approval itself, the items held over at the time of approval, and the items deliberately left undecided. The third is the one that gets skipped, and if it is not written down, “we had decided that, hadn’t we” arrives later.
And when a discussion opens after a Gate has been passed, do not call it rework. Once the structure is agreed, styling questions naturally follow, and those belong to the Production Harnesses, not to a rollback. Knowing what is genuine rework and what is simply the next step also depends on having an Architecture.
Writing the definition and keeping the working parts are two different things
Discussions about architecture usually begin with definitions. What roles to place, in what order, who approves. It is an enjoyable discussion and it produces a tidy document.
But what decides whether AI use takes root in an organization comes after that. The thinking gets written up and shared, while the checks and scripts that actually run stay in the environment of whoever built them. In that state, the mechanism ends the day that person moves on. This kind of loss — the document survives, the capability is not reproducible — happens more easily than people expect.
So we separate definition from implementation and put the implementation under version control. Outputs and their evidence are committed. Human Gate decisions are written with a date and a named decision-maker, and a tag is applied when something is settled. The client-specific and Neuromagic-common halves of the Standards / Knowledge Layer are separated as folders rather than as a stated principle, in a form that breaks if they are mixed. When separation is a working mechanism rather than a claim in a document, the boundary is still visible later.
The moat is in Experience Intelligence
Delivering the same scope faster and cheaper with AI is now the market standard, not a competitive advantage. Keeping old processes and old prices is not an option, and a company that does not notice will gradually stop being selected.
So what does differentiate? How well you can design, for each project, the appropriate scope, process, quality, schedule and budget allocation. And whether those design judgments come back to the organization every time a project runs. Which project conditions made which Harnesses effective, where placing a Human Gate stabilized the outcome, what quality, time, cost and risk actually measured, which failure patterns appeared. The circuit that accumulates this and feeds it back into how the next project is configured is what we call Experience Intelligence. That circuit is the substance of an Agent Harness Architecture, and it is the part that is hard to copy.
To summarize. What turns AI agents into results is not model performance, not prompt craft, and not building a single Harness. It is placing Harnesses both vertically and horizontally, connecting them through a Standards / Knowledge Layer and a structural source of truth, checking at Quality Gates, keeping Human Gates where people decide, and returning what is learned to Experience Intelligence. And none of it has to be complete at once. Take one module or one project, define its Input, Output, Questions and Human Gates, and run it. That is the minimum implementation unit, and the first step.
Where do the results of AI accumulate in your organization? If the honest answer is that they stay in individual hands, it may be time to build the design that connects your containers, before adding another one. We work on both the design of this Agent Harness Architecture and its operation in live projects. If you would like to think through where to start turning your own AI use into a system, please get in touch.
Feel Free to Contact Us
If you have any questions about the article or would like to discuss what these topics mean for your organization, please don’t hesitate to get in touch.

Motoharu Kuroi, President & CEO
Motoharu Kuroi graduated from the College of Liberal Arts at International Christian University and earned an MBA from the Kenichi Ohmae Graduate School of Business. Before founding Neuromagic in 1994, he spent seven years working in planning, production, and direction at an independent event producer’s office. In the early 1990s, he was involved in numerous projects that incorporated multimedia into event production and staging. Since the commercialization of the internet, he has led numerous digital and experience-related projects at Neuromagic.

