Some of the most misleading moments in AI are the impressive ones. A model writes a surprisingly good page, researches a difficult question, operates a browser, produces software or completes a workflow that would have required a person only a short time ago. The natural reaction is to think that the hard problem has been solved and that what remains is optimisation: a better prompt, more context or a stronger model. I used to lean much more heavily toward that view, but building with AI has gradually changed my mind.
A demonstration proves that a capability is possible. It tells you much less about whether that capability can become part of an organisation. The difference becomes visible when you ask the AI to do the useful thing again, then under slightly different conditions, and eventually with permission to affect something that matters. At that point the questions change. How do we know whether the result is good? What does the system remember? What authority does it actually have? What happens when it fails? Can we understand what happened afterwards? Can another model or runtime replace the current one without rebuilding everything around it? Who remains accountable when an output becomes a consequence?
The demo has not stopped being impressive. It has revealed the next problem.
From adding AI to designing for it
Most organisations are still at an earlier point in this transition. They are adding AI to software, processes and organisations that were designed before AI could meaningfully participate in the work. That is useful, may create enormous value, and for many organisations may be enough for quite some time. But I mean something different when I talk about AI-native systems.
An AI-enabled system adds intelligence to an existing design. An AI-native system starts from the assumption that intelligence can participate in the system itself: interpreting context, reasoning about situations, choosing among possible actions and adapting to circumstances that were not completely specified in advance. That changes the engineering question. AI experimentation asks whether AI can do something. AI-native engineering asks what makes that capability repeatable, coherent, measurable, governable and improvable once it begins to matter.
That second question can lead to an equally dangerous mistake. Once builders discover everything missing around a raw AI capability, the instinct is to build machinery around every weakness: another wrapper, another orchestration layer, a specialised retrieval system, another mechanism compensating for unreliable behaviour. I have become much more cautious about that instinct because AI capability is moving unusually quickly. Problems that genuinely required custom engineering a year ago can increasingly be handled by the model, runtime or open system itself. Machinery built to compensate for a temporary weakness can survive long after the weakness disappears, leaving an organisation carrying yesterday's assumptions inside tomorrow's technology.
The lesson after the demo is therefore not simply to build more machinery. It is to become more selective about what deserves to exist around the capability. My current working principle is: use native capability first, govern it, and replace selectively where evidence shows a real institutional gap.
Governance without suffocation
Governance is another word that can make this sound more complicated than it needs to be. I do not mean putting a committee in front of every AI action. I mean deciding what the system may do and access, where its authority ends, when a human must become involved, how we can reconstruct what happened and who remains accountable when something goes wrong. A model can become much smarter without gaining the right to approve a payment. A runtime can become much more reliable without deciding which organisational decisions require human judgement. Intelligence and authority are different things.
That distinction becomes increasingly important as AI moves from producing answers to participating in consequential work. The governance burden should therefore rise with consequence, persistence, autonomy and scale. A system drafting an internal note should not be governed like a system independently changing a customer's account. At the same time, good governance should not merely constrain capability; it should make useful capability safely usable.
That is harder than simply adding more controls. One reason AI is valuable is that it can reason through situations we did not enumerate in advance. If we respond by forcing every action through rigid workflows and approval gates, we risk rebuilding the limitations of traditional enterprise software around a technology whose value comes partly from adaptability. The objective cannot be maximum formalisation. It has to be enough structure for the consequence involved.
What should actually remain?
There is another reason I have become cautious about building too much around AI: the platforms themselves are moving. Identity, memory, evaluation, permissions, observability, tool use and orchestration are increasingly becoming capabilities of models, runtimes and cloud platforms. Some of the infrastructure organisations are building around AI today may eventually become ordinary platform capability, and I think that possibility should influence architecture now.
It does not mean organisations surrender responsibility to their providers. A platform may provide the mechanism for permissions, but it cannot decide what your organisation should legitimately permit. It may provide an audit trail, but it cannot decide what your organisation is accountable for. It may provide increasingly capable intelligence, but it cannot automatically determine which decisions your institution believes should remain human. The technical machinery may therefore become thinner even as the institutional questions become more important.
The durable layer around AI may not be the largest possible collection of wrappers, agents, orchestration frameworks and proprietary infrastructure. It may be the smallest sufficient layer that preserves what the organisation itself must continue to own: authority, accountability, continuity, judgement and the ability to change the underlying capability when circumstances change. I use “smallest sufficient” deliberately. A thin interface is not the same thing as a weak control boundary, and abstraction does not make the systems underneath it disappear.
The architecture of AI-native organisations remains unsettled, so any confident claim that the final form has already been found should be treated with suspicion. What feels more durable is a discipline: make experimentation cheap; make institutionalisation deliberate. Let intelligence explore where exploration is safe. Add structure as activity becomes persistent, delegated and consequential. Preserve human judgement where law, ethics, legitimacy or irreversible consequence require it. Use native capability first, and build around it only when evidence shows that the institution needs something more.
The successful demo is not the end of the hard problem. It is where the real shape of the problem becomes visible. A demo proves that AI can do something; an institution must decide what it may do, what it must remember, when it must stop and who remains accountable when the result matters. The question is no longer simply whether AI can perform a task, but what deserves to exist around that capability once the task becomes consequential, and what should be left out.
The future may belong not to the organisations that build the most around AI, but to those that learn exactly what must remain.
PUBLIC CLASSIFICATION: GREEN
CANONICAL SOURCE: ESSAY-001 · v0.5