Cian-OS Playbooks AI-native engineering
What Actually Changes in an AI-Native Engineering Team
In short
In an AI-native team, generating code stops being the bottleneck; the effort moves to defining small tasks, preparing context, setting boundaries for agents, and reviewing and verifying every change. What doesn’t change is accountability: AI executes, and people answer for the result.
Most of what gets published about engineering with AI is about speed: how much faster code gets written. To us, that’s the least interesting part. What really changes in an AI-native team is where human work concentrates: generating code stopped being the bottleneck, and the effort moves toward deciding what to build, preparing context, setting boundaries and verifying. This guide describes those changes in day-to-day work — and what doesn’t change.
What doesn’t change
We’ll start here, because it’s what gets lost most often:
- Accountability stays human. An agent can propose, write and test code. Whoever merges it answers for it.
- Understanding the problem is still the first job. AI speeds up executing a task; it doesn’t decide whether it was the right task.
- Architecture still gets decided. Generating quickly on top of an architecture nobody thought through just produces debt faster.
In Cian-OS we put it as a principle: AI executes; humans remain accountable.
What changes, in nine practices
1. Task decomposition
Tasks get smaller and more precise. Not because AI can’t produce a lot of code, but because a person has to be able to review the whole result. “Build the billing module” produces something nobody can review; “add tax-ID validation to the customer form, with its tests” produces something you can verify in minutes.
2. Context preparation
An agent only knows what it’s given. Project rules, conventions, architecture decisions, and the commands to test and deploy move out of someone’s head and into the repository, in writing. In Cian-OS we treat this as infrastructure: context is written, versioned and reviewed, just like code. A valuable side effect: that same context serves every new person on the team.
3. Agent and tool boundaries
You have to decide — in writing — what an agent may do without asking and what it may not. Reading code, proposing changes and running tests usually sit inside the line; deploying, touching production or deleting data sit outside it. Agents work without secrets or real customer data in their context.
4. Review
Code review stops being a formality and becomes the main control point. Nothing merges unless a person who understands the change has reviewed it, no matter who — or what — wrote it. The volume of changes goes up; the review standard can’t go down.
5. Verification
Every gain in generation speed widens the gap between what gets produced and what has been checked. If verification doesn’t scale at the same pace, that gap fills with errors nobody saw. That’s why, in the way we work, verification isn’t a phase at the end: automated tests, static analysis, dependency checks, security checks, human review and acceptance validation run continuously.
6. Tests
AI writes tests easily, and that has a catch: a test that never fails proves nothing. Someone has to check that generated tests fail when they should, and review them as carefully as the code they protect.
7. Security
New risks appear and some old ones get worse. An agent can propose a dependency nobody evaluated, or reproduce an insecure pattern with complete confidence. And if the product uses language models, their inputs must be treated as untrusted: prompt injection tops OWASP’s list of risks for applications built on language models. Changes to authentication, permissions and sensitive data get stricter review.
8. Human judgment
When execution is cheap, value concentrates in deciding: which problem is worth solving, what gets built and what doesn’t, which architectural trade-off will age well, when something is ready for production. Those decisions aren’t delegated to a model, because nobody can hold a model accountable.
9. Delivery governance
Measuring how much code AI produces measures the wrong thing. What’s useful is measuring how much verified software reaches production and how much of it holds up: indicators like deployment frequency, the time it takes a change to reach production, the rate of changes that fail and the time to recover — the ones DORA proposes — say more than any count of lines or tasks.
Why this Playbook doesn’t promise productivity
We don’t include productivity multipliers. The figures in circulation measure different things — lines, tasks, time to a first version — in different contexts, and almost never measure what matters: verified software, in production, that the team can keep maintaining. If AI adoption is working, it shows up in those delivery indicators. If it only shows up in code volume, it isn’t working yet.
An illustrative example: one task, end to end
A four-person team (an illustrative example) needs the ordering system to reject discounts above each customer’s authorized limit:
- Task: one person scopes it — the rule, where it applies, what message the user sees, and which tests must pass.
- Context: the agent works with the repository’s conventions and the already-documented decision that pricing rules live on the server.
- Boundaries: the agent can change code and run tests in its environment; it can’t deploy or see customer data.
- Generation: the agent proposes the change and its tests.
- Verification: tests run on every change; an engineer confirms they fail if the validation is removed.
- Review: an engineer who knows the pricing module reviews the whole change; because it touches money rules, it gets stricter review.
- Delivery: the change reaches production through the usual deployment, with rollback available. A named owner answers for it.
Nothing in this flow depends on a particular tool. All of it depends on someone having designed it.
Common mistakes
- Measuring generated code instead of verified software in production.
- Handing an agent huge tasks nobody can review in full.
- Leaving context in people’s heads and expecting agents to guess it.
- Giving agents access to secrets or production “so they’re more useful.”
- Lowering the review bar because the volume of changes went up.
- Trusting generated tests without checking that they fail when they should.
- Choosing tools before designing the workflow.
Practical tool · Checklist
AI-native engineering workflow checklist
A way to review how a team that uses AI to produce software actually works today. It isn’t a tool list: every question holds for any model, agent or editor. Mark each one Yes, Partly or No.
A · Before asking for code
| Question | What “in place” looks like | Status (Yes · Partly · No) |
|---|---|---|
| Is each task small enough for one person to review the whole result? | Tasks with a verifiable result — not “build the module.” | |
| Does the context live in the repository? | Rules, conventions, architecture decisions and commands written down and versioned, not in someone’s head. | |
| Is “done” defined before anything is generated? | Acceptance criteria and expected tests written before work starts. |
B · Agent and tool boundaries
| Question | What “in place” looks like | Status (Yes · Partly · No) |
|---|---|---|
| Is it written down what an agent may do without asking? | Allowed actions (read, propose, run tests) and forbidden ones (deploy, touch production, delete data). | |
| Do agents work without secrets or production data? | Credentials kept out of the context; test or anonymized data. | |
| Can you tell which changes an agent proposed? | Traceability in review or in the change history. |
C · Review and verification
| Question | What “in place” looks like | Status (Yes · Partly · No) |
|---|---|---|
| Is every change reviewed by a person who understands it? | Nothing merges without human review, no matter who — or what — produced it. | |
| Do automated tests run on every change? | A suite that fails when something breaks, reviewed like any other code. | |
| Do AI-generated tests actually test something? | Someone has checked that they fail when they should. | |
| Do static analysis and dependency checks run before merging? | Including new dependencies an agent proposed adding. |
D · Security
| Question | What “in place” looks like | Status (Yes · Partly · No) |
|---|---|---|
| Do changes to authentication, permissions or sensitive data get stricter review? | A list of sensitive areas with mandatory review by someone with security judgment. | |
| If your product uses language models, does it treat their inputs as untrusted? | Defenses against prompt injection, and a model with no permissions it doesn’t need. |
E · Accountability and governance
| Question | What “in place” looks like | Status (Yes · Partly · No) |
|---|---|---|
| Does every delivery have an accountable person? | Accountability isn’t delegated to a tool. | |
| Do you measure verified software in production, not generated code? | Delivery indicators — frequency, changes that fail, time to recover — instead of lines or tasks produced. | |
| Does the system stay understandable to someone who didn’t write it? | Another engineer, or an agent with the repository’s context, can pick up the work. |
F · How to read the result
- “No” in A
- AI will quickly produce something that’s hard to review. Start with context and task size.
- “No” in B or D
- It’s a security risk before it’s a productivity one. Fix it before giving agents more autonomy.
- “No” in C
- The gap between what’s generated and what’s verified is growing — and more speed widens it.
- “No” in E
- Nobody is accountable for the result. That’s the first fix.
How it connects to Cian-OS
These practices are the core of Build — people and AI agents executing within explicit context and boundaries — and of Verify, where verification is a system of its own rather than a phase at the end. It’s how our Product Engineering teams and Engineering Pods work.
Sources
- OWASP Top 10 for LLM Applications — LLM01:2025 Prompt Injection — Prompt injection as the first entry on OWASP’s list of risks for applications built on language models.
- DORA — Software delivery performance metrics — Software delivery indicators — deployment frequency, change lead time, change fail rate, recovery time — rather than code volume.
Engineering Pods
If you want to work this way without building the system from scratch: an Engineering Pod adds engineers to your roadmap who already work with versioned context, human review of every change and continuous verification.