By

Governing AI-Assisted Software Engineering: Why “Vibe Coding” Needs Guardrails, Not a Ban

AI is no longer a nice-to-have in software engineering. It’s embedded everywhere, from inline code completion to autonomous agents that plan entire pieces of work and open their own pull requests. The most permissive end of this spectrum is what Andrej Karpathy called “vibe coding” — leaning fully into the suggestions an AI assistant gives you and letting go of the need to read every line yourself.

That posture might be fine for a weekend side project. It is not acceptable in an enterprise setting, where AI-assisted development touches security, intellectual property, quality, and regulatory compliance all at once. The organisation needs a way to capture the real productivity gains of AI-assisted engineering without inheriting its risks unmanaged. That’s what a governance framework for AI-assisted software engineering is for.

Here’s a walkthrough of what such a framework covers and why each piece matters.

It’s Not Just About Code

The instinct is to think of AI governance as “reviewing AI-written code.” In practice, AI assistance reaches much further than source code: it shapes requirements, architecture, tests, documentation, infrastructure-as-code, deployment pipelines, and even the technical decisions teams make along the way.

That last category is easy to underestimate. Artefacts that aren’t executable — requirements, architecture, decisions — often get less scrutiny than code, yet they can do more damage, because everything downstream gets built on top of them. A flawed architectural choice or a hallucinated requirement doesn’t show up as a failing test; it shows up months later as a rebuild.

Risk Should Set the Bar, Not the Tool

Not every piece of work carries the same stakes, and a governance framework worth having scales its controls accordingly. A useful way to think about this is a simple risk tiering model:

  • Low — internal utilities, scripts, throwaway tooling. Author self-review is enough, and the AI can generate fairly freely.
  • Medium — internal business applications. One independent, qualified reviewer signs off before merge.
  • High — customer-facing applications. Independent review plus a security review; no unattended generation.
  • Critical — financial, healthcare, or safety systems. Two independent reviewers (one from security), full regulatory validation, and AI used as an assist only — never autonomously.

The point isn’t to slow everyone down uniformly. It’s to make sure the depth of oversight actually matches what’s at risk if something goes wrong.

Autonomy Is a Separate Dial From Risk

Risk tier tells you how critical a system is. Autonomy mode tells you how much of the work the AI is doing without a human directing each step. These are independent variables, and conflating them is a common mistake.

A framework can define three autonomy modes, in increasing order of delegated authority:

  1. Human-assisted development — the person directs the work step by step; the assistant completes or suggests within that. This is the default for most day-to-day coding.
  2. Human-supervised autonomous development — the human sets an objective, the agent plans and executes multi-step work (editing files, opening pull requests), and a qualified person reviews before anything merges. The agent proposes; a human disposes.
  3. Fully autonomous experimentation — the agent runs end to end without step-by-step approval, but only inside an isolated sandbox with zero access to production systems, data, secrets, or customer-facing environments.

Cross-reference these with risk tier and you get a gating matrix: an agent working with minimal supervision on an internal utility is low overall risk; that same agent working the same way against a payments system is not. Fully autonomous experimentation should never be permitted against production, full stop, regardless of tier.

Accountability Has to Land on a Named Person

Perhaps the most important — and most often fudged — principle is this: a tool is never accountable. A person is. And that person needs to be named, not implied.

The cleanest way to resolve this is to tie accountability to the decision to merge, not to the act of prompting. Whoever reviews, approves, and merges a change is the accountable owner, because that’s the last human decision before something ships. A prompt author who never merges is responsible for the quality of what they proposed, but they aren’t the accountable owner of what someone else chose to ship. When an autonomous agent opens a pull request, the person who approves and merges it carries that accountability — not the agent, not the prompt author.

This matters most when something breaks. If a production incident traces back to AI-generated output, the fact that “the AI wrote it” is never a defence and never transfers responsibility elsewhere.

Architecture Is Where AI Does Its Quietest Damage

Code review catches bad code. It’s much worse at catching bad design, because each pull request can look individually sound while the system as a whole drifts. This is exactly the failure mode AI amplifies: it will happily produce duplicated architectures, apply patterns that don’t fit the context, over-engineer simple services, and introduce abstractions that add complexity without adding value — and every one of those changes can sail through a normal code review.

The fix is to require an Architecture Decision Record (ADR) for any architecturally significant change: new services, new datastores, new external dependencies, changes to cross-cutting patterns like auth or messaging, deviations from approved designs, or new frameworks entering the codebase. Where an AI proposed the design, the ADR should say so explicitly, and the human author should confirm they actually understand and endorse it — not just that it compiled.

Data Handling Needs Explicit Rules, Not Good Intentions

Prompts are a data-exfiltration surface whether anyone thinks of them that way or not. A workable framework classifies data and sets hard rules:

  • Public data — no restriction.
  • Internal data — only with tools contractually guaranteed not to train on or retain it.
  • Confidential or proprietary source — only with enterprise-contracted tools offering exclusion guarantees, and only with team-lead approval.
  • Personal data (GDPR-relevant) — prohibited by default, unless there’s a documented lawful basis and DPO sign-off.
  • Secrets and credentials — never, under any circumstances.

Models Themselves Are Governed Components

One detail that’s easy to miss: the model behind an approved tool can change without anyone asking. Vendors ship new versions, sometimes silently, and a model’s behaviour can shift meaningfully between them. Treating “the tool” as approved isn’t enough — the specific model version needs to be the unit of approval, time-bound and subject to re-validation. When a vendor deprecates or materially changes an approved model, that approval should lapse for high-stakes work until it’s been reassessed. Periodic sampling of generated output for hallucination rate and behavioural drift closes the loop.

Why Bother With All This?

Because the alternative isn’t “no governance” — it’s invisible governance, where risk accumulates quietly until an incident forces the conversation anyway. A clear framework does the opposite: it lets teams move fast on low-stakes work with light-touch review, while making sure the controls tighten automatically as the stakes rise. Vibe coding has its place. It just isn’t in charge of anything that can hurt the business, its customers, or its data if it goes wrong.

Leave a Reply

About the blog

RAW is a WordPress blog theme design inspired by the Brutalist concepts from the homonymous Architectural movement.

Get updated

Subscribe to our newsletter and receive our very latest news.

← Back

Thank you for your response. ✨

Discover more from The Golden Hour

Subscribe now to keep reading and get access to the full archive.

Continue reading