# Overview

Welcome to the Kerno documentation.

### What is Kerno?

Kerno learns how your product works, maps the user flows through your app, and records how they behave in tests that run against your real stack. Every time your code changes, Kerno validates the change inside your coding agent's loop and shows what behavior moved.

Kerno runs locally. Your coding agent drives it over MCP, and the portal reports across your team.

### Start here

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><i class="fa-rocket-launch" style="color:$info;">:rocket-launch:</i> <strong>Get started</strong></td><td>Install Kerno and connect it to your coding agent.</td><td><a href="/docs/getting-started/quickstart">Quickstart</a></td></tr><tr><td><i class="fa-lightbulb" style="color:$info;">:lightbulb:</i> <strong>How Kerno works</strong></td><td>Understand how Kerno works and where it fits in your agentic code workflow</td><td><a href="/docs/core-concepts/how-kerno-works">How Kerno Works</a></td></tr><tr><td><i class="fa-layer-group" style="color:$info;">:layer-group:</i> <strong>Supported technologies</strong></td><td>Check your language, framework, and datastore.</td><td><a href="/docs/references/supported-technologies">Supported Technologies</a></td></tr><tr><td><i class="fa-arrows-turn-to-dots" style="color:$info;">:arrows-turn-to-dots:</i> <strong>Changelog</strong></td><td>Stay up to date with the latest feature and improvements</td><td><a href="/docs/getting-started/changelog">Changelog</a></td></tr></tbody></table>

### Key benefits

**For Developers**: Ship faster with AI, without breaking your system.

* Find, fix and verify inside the session, not minutes or hours later in CI
* Validate every AI-generated change against your real stack before raising a PR
* Avoid lengthy reworks from issues discovered after you've already pushed
* Never manually write, update, or maintain your tests

**For Teams**: Maintain quality and consistency as your team ships with AI

* Every developer ships to the same standard of reliability
* Adopt AI across your team without increasing production risk
* Entry point coverage that stays current as your codebase evolves, with no manual upkeep
* Increase AI usage across your team without adding QA overhead

### Resources

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><i class="fa-comment-question" style="color:$info;">:comment-question:</i> <strong>Support</strong></td><td>Get help from the Kerno team.</td><td><a href="https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ">https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ</a></td></tr><tr><td><i class="fa-lightbulb" style="color:$info;">:lightbulb:</i> <strong>FAQs</strong></td><td>Get answers to common questions.</td><td><a href="/docs/getting-started/faqs">FAQs</a></td></tr><tr><td><i class="fa-shield-check" style="color:$info;">:shield-check:</i> <strong>Security</strong></td><td>Learn how Kerno handles your code and data.</td><td><a href="/docs/references/security-and-privacy">Security &amp; Privacy</a></td></tr></tbody></table>


# Quickstart

Install Kerno and connect it to your AI coding agent.

### Prerequisites

* **Docker** (or Podman): must be installed and running
* **Node.js 18+**: required to run the Kerno CLI
* **Git**: your project must be a git repository

### Set up Kerno with your agent

Kerno is driven by your coding agent, so let your agent install it. Paste this into any coding agent and it will install the CLI, bind Kerno to your repository, register the MCP server, and kick off onboarding

```
Set up Kerno for this repo: install the CLI, bind it to this workspace, register MCP with my coding tool, then trigger workspace analysis and list the available apps for me to choose from. Follow the Kerno MCP's instructions throughout. Don't run any tests, and don't set up the environment yet.

Read rather than improvise. kerno init prints what you need, and once MCP is live every Kerno tool carries a What/When description saying what to call next. Follow those. The commands below are the happy path; if kerno init prints something different, its output wins.

Prerequisites that must already exist. If any of these is missing, stop and tell me: Node.js 18+, npm, Docker or Podman running, and a git repo.

  npm install -g @kerno/cli
  kerno login
  kerno init -w "<absolute path to this repo>"

Don't run kerno login yourself. It renders an interactive terminal UI and dies immediately in an agent shell with "Raw mode is not supported on the current process.stdin", so the browser never opens. Instead, tell me to run ! kerno login in the prompt myself, then pause and wait for me to confirm I've logged in before continuing. If I'm working over SSH or in a container with no browser, tell me I can log in with kerno login --api-key instead, or by setting KERNO_API_KEY.

Run kerno init -w even if something seems to be listening already, never two at once, then use the workspace, port and registration snippet it prints verbatim. The port can change when its usual one is taken, so never reuse one from docs, an old config, or memory. In an agent shell, kerno init prints the port and registration snippet as plain text. If they are missing, read the port from ~/.kerno/mcp.port and build the endpoint as http://localhost:N/mcp. kerno doctor --clean fixes an orphan or inconsistent agent.

If kerno init reports a problem with the organization or its key, stop and ask me which organization to use, then re-run kerno init -w "<path>" --org <name or id>. Never edit ~/.kerno/state.json by hand.

Register with the host you are running in, and ask whether I want a different one. Register at one scope only, project or user, and merge rather than overwrite other servers. Then allowlist Kerno before calling any tool. A single task runs many calls, including repeated status and job polling, so without this I am clicking approve every few seconds. Explain that to me, show me the change, and apply it once I accept. If it's already present, tell me and move on. Claude Code: "mcp__kerno__*" in permissions.allow in .claude/settings.json or ~/.claude/settings.json. Codex: default_tools_approval_mode = "approve" under [mcp_servers.kerno] in ~/.codex/config.toml. Cursor: no file, tell me to set Run Mode to Auto-review.

Registering the server does not load its tools into a session that was already running. Before you try to call anything, expect the kerno_* tools to be absent, and tell me to reconnect: in Claude Code run /mcp, select the server, and reconnect, or restart the tool. A "Connected" line from claude mcp list is not proof your current session can call the tools, because that check opens its own session.

Verify with kerno_get_applications; that succeeding means connected.

If a stop, restart or workspace switch moves Kerno to a different port, take the new endpoint URL from kerno init's output and reconnect with it.

Once connected, let Kerno finish analyzing the workspace. The first analysis can run for several minutes and outlast a single tool call, which looks like a hang but isn't. Calling again with the same arguments is safe and picks up the same run. Don't sync the workspace or clear the cache to unstick it, because that restarts the analysis from scratch.

Then list the apps it found and present them to me so I can choose which one to test. If there is more than one, mark which one looks like the backend, and once I pick, name that one app in every later Kerno call. Calls without an app analyze every app in the repo, and each app gets its own sandbox container. Stop there and follow the Kerno MCP's tool guidance for what comes next, including the usage note it asks you to add to this repo's agent instruction files.
```

### Install manually

#### 1. Install the CLI

```bash
npm install -g @kerno/cli
kerno login
```

`kerno login` opens a browser to authenticate you. On a machine with no browser, such as over SSH or in a container, run `kerno login --api-key <key>` or set `KERNO_API_KEY`. Once signed in, return to your terminal and point Kerno at your repository:

```bash
kerno init -w /absolute/path/to/your/repo
```

Or run `kerno init` from inside the project directory. On first run this downloads the Kerno agent, which takes a moment. It then binds the agent to that workspace and prints the MCP server URL, along with a ready-made registration command for Claude Code and a config snippet for Cursor.

Each repository runs under its own Kerno organization. The first time you run `kerno init` in a repository, it asks which organization to use there and remembers your answer. Add `--org <name or id>` to choose it up front, for example when no terminal is available to ask in.

{% hint style="info" %}
`kerno init` binds the agent to **one repository root**. The agent then serves that root and every directory nested under it as separate workspaces, so a monorepo's applications are all reachable from the one agent, selected per call by `workspace_path`. To point it at a repository **outside** that root, prefer `kerno stop` followed by `kerno init -w <other-path>`, which cancels the current workspace's in-flight work before shutting down. In a terminal, `kerno init -w <other-path>` alone asks whether to stop the running agent and start there. From an agent shell, add `--force-switch` to switch in one step, though it stops the old agent while its work is still running.

If switching workspaces or restarting the agent moves Kerno to a different port, update your coding tool with the URL `kerno init` prints.
{% endhint %}

#### 2. Connect MCP

Use the command or config snippet that `kerno init` printed. It already contains the correct URL, and any MCP-compatible coding tool can connect to it.

{% hint style="info" %}
Kerno keeps the same MCP port across restarts, 8086 by default, so your config keeps working. If another program holds that port, Kerno picks a free one, and `kerno init` prints the new URL.
{% endhint %}

Then allow Kerno's tools, or you will approve every call by hand. A single task runs many Kerno calls, including repeated status and job polling.

* **Claude Code:** add `"mcp__kerno__*"` to `permissions.allow` in `.claude/settings.json` for this project, or `~/.claude/settings.json` for every project.
* **Codex:** add `default_tools_approval_mode = "approve"` under `[mcp_servers.kerno]` in `~/.codex/config.toml`.
* **Cursor:** set Run Mode to **Auto-review**, which lets allowlisted MCP tools run without prompting.

### Setup Kerno for your project

Once MCP is connected, ask your agent to take it from here:

```
Use the Kerno MCP to set up Kerno for this project.
```

Your agent checks that Kerno is reachable, asks which applications are in your repository, and picks one to work on. It saves the URL your application runs on and, optionally, credentials for your database and other dependencies, asking you for whatever it cannot find in your repository. Once Kerno reports the environment ready, your agent calibrates Kerno to your repository and asks you to confirm what it found. It then maps your key [user flows](/docs/core-concepts/user-flows) with you. See [Calibration](/docs/core-concepts/environment-setup#calibration). The setup page in the Portal shows each step as it completes.

### What's next

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Test a User Flow</strong></td><td>Learn how to test a key journey through your app end to end</td><td><a href="/docs/guides/test-a-user-flow">How to Test a User Flow</a></td></tr><tr><td><strong>Configure the Kerno Test Environment</strong></td><td>Learn how to connect Kerno to your running application</td><td><a href="/docs/guides/start-the-environment">How to Configure the Kerno Test Environment</a></td></tr><tr><td><strong>Create Baseline Tests for your Entry Points</strong></td><td>Learn how to generate baseline tests for an entry point</td><td><a href="/docs/guides/capture-a-baseline">How to Create Baseline Tests for your Entry Points</a></td></tr></tbody></table>


# Uninstall Kerno

Remove the MCP registration from your coding tool, stop the local agent, and clean up the installed binaries and the CLI.

### Uninstall with your agent

Paste this into the coding agent Kerno is connected to.

```
Uninstall Kerno from this machine. Show me each change before you make it.

Remove the kerno entry from my coding tool's MCP config, and the allowlist rule that went with it:
"mcp__kerno__*" in permissions.allow for Claude Code, or the
[mcp_servers.kerno] block for Codex. Ask me which host if it isn't obvious.

  kerno logout                 clears the stored login
  kerno uninstall              stops the agent and removes the downloaded runtime
  npm uninstall -g @kerno/cli

Then tell me what is still on disk and leave all of it alone unless I ask you to delete it.
```

### Uninstall manually

#### 1. Remove the MCP entry

Edit your coding tool's MCP config file and delete the `kerno` server block.

#### 2. Stop the agent and remove its binaries

```bash
kerno uninstall
```

This stops the running agent and removes the downloaded agent runtime from `~/.kerno/assets/`.

#### 3. Uninstall the CLI

```bash
npm uninstall -g @kerno/cli
```

### What Kerno leaves behind

`kerno uninstall` clears `~/.kerno/assets/` and nothing else. The following stay on your machine until you remove them yourself:

| Location                                                                             | What it holds                                                          |
| ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------- |
| `~/.kerno/workspaces/`                                                               | Per-workspace snapshots, indexed data, prompt cache, and logs          |
| `~/.kerno/state.json`                                                                | Which organization each repository uses, and where Kerno is turned off |
| `~/.kerno/secrets.json`                                                              | Your Kerno login credentials, when the OS keychain can't hold them     |
| `~/.kerno/agent/`, `agent.stdout.log`, `cli.installed`, `cli-update-check.json`      | Agent working files, its output log, and CLI install and update checks |
| `.kerno/` in each repository                                                         | Your scenarios, baselines, and workspace config                        |
| `.agents/skills/kerno-*` and `.claude/skills/kerno-*` in each repository             | The Kerno skills for your coding agent, and links to them              |
| `.git/info/exclude` in each repository                                               | The lines that keep those skills out of git                            |
| OS keychain                                                                          | Your Kerno login credentials                                           |
| Docker containers named `kerno-ts-sandbox-*`, and the `kernoio/ts-sandbox` image     | The sandbox your tests run in                                          |
| `kerno/sandbox/` in your system temp directory, and `/tmp/kerno/scenario-workspace/` | Sandbox config and scratch copies of your tests                        |

Run `kerno logout` before uninstalling to clear your stored credentials. Delete `~/.kerno/` if you want the local data gone, and remove the sandbox containers and image with Docker. Keep each repository's `.kerno/` directory if you may reinstall later, since it holds the test suites Kerno generated for that codebase.

### Resources

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><i class="fa-comment-question" style="color:$info;">:comment-question:</i> <strong>Support</strong></td><td>Get help from the Kerno team.</td><td><a href="https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ">https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ</a></td></tr><tr><td><i class="fa-lightbulb" style="color:$info;">:lightbulb:</i> <strong>FAQs</strong></td><td>Get answers to common questions.</td><td><a href="/docs/getting-started/faqs">FAQs</a></td></tr><tr><td><i class="fa-shield-check" style="color:$info;">:shield-check:</i> <strong>Security</strong></td><td>Learn how Kerno handles your code and data.</td><td><a href="/docs/references/security-and-privacy">Security &amp; Privacy</a></td></tr></tbody></table>


# FAQs

A collection of answers to frequently asked questions about Kerno

<details>

<summary>What is Kerno?</summary>

Kerno learns how your product works, maps the user flows through your app, and records how they behave in tests that run against your real stack. Every time your code changes, Kerno validates the change inside your coding agent's loop and shows what behavior moved.

Kerno runs every change against your real stack and validates it across functional behavior, error handling, edge cases, security, auth, and response content.

</details>

<details>

<summary>What makes Kerno different from other tools</summary>

Most code tools either statically analyze your code or run unit tests in isolation against mocks. Kerno does neither.

**It knows the exact blast radius of every change.** Kerno indexes your code deterministically, so it can tell you precisely which entry points a change affects, following the call graph rather than returning approximate matches from vector embeddings.

**It tests against your real stack.** Scenarios run against your actual running application and its real databases and queues, not mocks.

**You don't write or maintain the tests.** Kerno generates them, and updates them when your entry points change on purpose. When it writes new tests for an entry point, it shows you the plan and waits for your approval first. Updates and batch runs go ahead unless you ask to review the plan.

</details>

<details>

<summary>Does Kerno support my backend stack?</summary>

Most likely, yes. Kerno detects routes in TypeScript, JavaScript, Python, Java, Kotlin, Scala, Go, Ruby, PHP, C#, Rust, and Swift, across the common web framework for each.

Your scenarios are always written in TypeScript and reach your application over HTTP, MCP, or the queue a background consumer reads, so the language your service is written in doesn't change how tests run.

Kerno can also connect directly to PostgreSQL, MySQL, MariaDB, MongoDB, Redis, Kafka, RabbitMQ, Amazon SQS, Azure Service Bus, DynamoDB, ClickHouse, Azure Storage and S3-compatible object storage, for setting up and verifying state your API can't express. It also reads Zitadel and Amazon Cognito to obtain test credentials.

The full matrix lives in [Supported technologies](/docs/references/supported-technologies).

</details>

<details>

<summary>Does Kerno store my code</summary>

No. Kerno never stores your source code on its servers. Relevant excerpts pass through Kerno's proxy to LLM providers (Anthropic and OpenAI) configured with zero-data-retention tiers, so nothing is persisted on the provider side either.

Note that the report of every test run your agent starts is uploaded as the run progresses, so your team can view it in the portal. That report includes the request and response payloads your application produced during the run. See [Security & Privacy](/docs/references/security-and-privacy).

</details>

<details>

<summary>Does Kerno use my code for training or fine tuning LLMs?</summary>

No. Kerno doesn't train models. The LLM providers Kerno uses run on zero-data-retention tiers, so your code isn't stored, logged, or used for training on their side.

</details>

<details>

<summary>Are there discounts available for startups?</summary>

Kerno offers free usage for qualified Open Source Projects and a significant discounts for pre-Series A startups. [Drop us a note](https://www.kerno.io/contact) to check your availability.

</details>

<details>

<summary>Can Kerno test an app that spans several repositories?</summary>

Kerno works within one repository at a time. That repository can hold several apps, such as the services of a monorepo, and Kerno tests each of them.

</details>

<details>

<summary>Does Kerno work with my coding agent?</summary>

Yes, if your agent supports MCP. Kerno exposes a set of MCP tools your agent calls to connect to your running app, generate tests, and validate changes. Claude Code, Cursor, Windsurf, and Codex are the verified setups.

See the [Quickstart](/docs/getting-started/quickstart) to install the CLI and register the MCP server with your agent.

</details>

<details>

<summary>What if I'm not using a coding agent?</summary>

Kerno runs through a coding agent that supports MCP, such as Claude Code, Cursor or Codex. The agent connects Kerno to your running app, generates tests for an entry point, and validates your changes. See the [Quickstart](/docs/getting-started/quickstart) to set one up. Once your tests are committed, the [CI action](/docs/guides/run-tests-in-ci) re-runs them on your pull requests without an agent.

</details>

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we’ll gladly help.
{% endhint %}


# Changelog

Stay up to date with the latest Kerno feature and improvements

### September 25, 2026

Kerno CLI 0.2.2 is on npm as [`@kerno/cli`](https://www.npmjs.com/package/@kerno/cli). Update with `npm install -g @kerno/cli`.

#### What's new

**Test a user flow end to end.** Ask your agent to test a journey through your app, such as a customer paying for a product. Kerno plans one test that walks the steps in order and checks the outcome at the end, and you approve the plan before it runs. After a code change, your agent can find the flows the change reaches and re-run them. See [How to Test a User Flow](/docs/guides/test-a-user-flow).

**Test MCP servers that require authentication.** Add your server's auth headers to `MCP_AUTH_HEADERS` in `.kerno/.env`, and Kerno sends them with every test call. See [Testing MCP servers](/docs/core-concepts/scenarios-and-baselines#testing-mcp-servers).

#### Improvements & Bug fixes

**Run reports show what MCP and background-job tests did.** A report now carries the MCP call the test made, or the job it dispatched, so you can check the result against it.

**Computed values are checked as values.** A count or a total your code computes is asserted at its exact value, so a wrong number fails the test.

**Test descriptions are written safely.** A description with a quote in it used to break the test file, and a test that had passed was then reported as failed.

**Your MCP credentials stay out of the logs.** A malformed `MCP_AUTH_HEADERS` value is logged by its error type only, so the token inside it is never written to a log.

**Kerno's MCP server accepts browser requests only from your own machine.** Before, any web page could reach it.

### September 22, 2026

Kerno CLI 0.2.1 is on npm as [`@kerno/cli`](https://www.npmjs.com/package/@kerno/cli). Update with `npm install -g @kerno/cli`.

#### What's new

**Critical entry points.** Mark an entry point as critical in the portal, and Kerno always tests it at high effort, whatever effort your agent asks for. See [Critical entry points](/docs/portal/coverage#critical-entry-points).

**A portal run report for every re-check.** When your agent re-runs your tests, each entry point gets its own run report in the portal, so a CI run links to a report for every entry point it checked.

#### Improvements & Bug fixes

**Leaked error details fail the test.** A response that exposes a stack trace, a file path or a database error fails its test and can't become the baseline. A test that passed on an error response must also check what the body carried.

**Review what Kerno concluded from its analysis.** Each analysis is split into records you can review with your agent, next to its lessons, and a wrong claim can be corrected on its own.

**Setup progress is live in the portal.** Calibration shows as running for as long as it runs, each step says when it finished, and the portal mirrors your agent's jobs as they happen.

**Test setup matches your app's sign-in method.** Kerno writes setup for Basic auth, API-key headers, Digest and DPoP, and gives a cookie-based app a cookie session.

**CI publishes what was really sent.** A run replayed in CI writes the real requests and responses next to its JUnit report.

**A batch of tests can be stopped.** Ask your agent to cancel a batch, as with any other run.

**Potential bugs point only at what the run observed.** When a run can't place the fault, it says so.

**More accurate Python route discovery.** An HTTP method Kerno can't establish is reported as unknown, a viewset's path comes from its own router, and a route inside an inline `include([...])` keeps its prefix.

**Session cookies are recognised under any name,** such as `connect.sid` or `next-auth.session-token`.

**Setup names the healthcheck error.** When the "Environment connected" step fails, it says what went wrong.

**An unreadable `.kerno/config.yaml` says where its copy went.**

**Each run counts once.** One agent run is one row in the portal's runs.

**An agent that goes silent shows as disconnected.**

### September 14, 2026

Kerno CLI 0.2.0 is on npm as [`@kerno/cli`](https://www.npmjs.com/package/@kerno/cli). Update with `npm install -g @kerno/cli`.

#### What's new

**Follow a run from the moment it starts.** `kerno_endpoint_test` hands back a portal link as soon as the run is accepted, the run report opens before the run body starts, and the run reports its current phase and the scenarios it plans to try while it is still going. A run in flight now shows what it is doing and what it will cover, rather than a spinner that resolves only at the end.

**Test your Graphile Worker background jobs.** Kerno discovers Graphile Worker tasks as consumer endpoints and plans them with a prompt written for that transport, so the jobs that run behind your API are tested next to your HTTP routes.

**What Kerno learns about your application is published into your repository.** Each record is a Markdown file with a YAML header under `.kerno/memory`, with a generated README beside them, so a record can be read, diffed and reviewed like any other file in your project. Two branches that learned about different endpoints now merge cleanly, where they used to collide on one JSON array.

**Run reports name the memory behind them.** When a plan was shaped by a lesson Kerno recalled, the report says which record contributed.

#### Improvements & Bug fixes

**Closed value sets are declared in the schema.** The MCP arguments with a fixed vocabulary, `effort`, `box_testing_strategy`, `tags`, `filter`, `verdict` and `reviewed_by`, are now schema enums. The accepted values are unchanged, and a value outside the set is refused at the protocol layer with the vocabulary named, where `"MEDIUM"` and `" agree "` used to be quietly repaired.

**Scenario planning reads the parameters your handler actually branches on.** Planning now works from the code paths rather than from parameter categories alone. A seven-parameter list endpoint used to get five scenarios covering a single parameter, and passed 5/5 on both the buggy and the fixed commit.

**Your instructions for a run outrank recalled memory.** A stored conclusion that is wrong about your application can be corrected by saying so at the run.

**Clearer tool descriptions.** Re-showing a finished run is separated from re-running it, `kerno_ignore_potential_bug` names all three fields it requires, and the guidance covers both ways to accept a flagged diff.

**Every organization-scoped read checks organization membership.** Previously these reads confirmed only that the caller was a valid Kerno user.

**`kerno_job` returns the structured content it promises.** It declared an output schema and returned no structured content, so a schema-checking host rejected the whole payload instead of showing the poll's result.

**A run parked at an approval gate says so.** An endpoint test waiting on approval used to report itself as `running` and tell the caller to check back for progress, advice a gate can never satisfy. It now reports `needs_user_feedback`, with the guidance that clears it.

**Concurrent runs keep their own activity logs.** Runs in flight shared one log, so every run collected every other run's task lines as if they were its own.

**Scenario and precondition resources are reachable.** They were listed but never resolved, because their paths are multi-segment and the templates only ever matched a single segment.

**Endpoint discovery reports from your workspace.** It was reporting from the snapshot clone, so every endpoints event was dropped for every repository and coverage read `0 / 0` on the repositories under the heaviest test. A repository already filed under a checkout path is now reconciled onto its forge identity and keeps its endpoints.

**`kerno_get_state` and `kerno_cancel` agree on a job id.** The two surfaces used to disagree about the same job at the same moment.

**Workspace listing and index status work with any number of workspaces.** Both failed outright unless exactly one workspace was registered.

**A config file Kerno cannot read is preserved.** It used to be rebuilt from scratch, which left one patched application in place of every application in the file, with nothing said.

**A `preconditions.ts` that compiles and always fails is rebuilt.** Staleness was judged by whether the script compiled, so one naming invented environment variables produced the same environment fault forever.

**A consumer task that cannot be dispatched is refused.** A crashed task changes nothing, so a scenario asserting that nothing changed used to pass on it.

**A failed authentication no longer discredits an unrelated lesson.** A 401 that a record cannot vouch for used to spend the confidence of a lesson that could never have produced it.

**Seeded rows respect your UNIQUE constraints.** Rows seeded by a generated scenario derive their UNIQUE constraints from the schema, so a second run, or two scenarios running at once, no longer fails in arrange on a constraint violation that says nothing about the endpoint.

**Credential-shaped material is scrubbed at every door into memory.**

**Generated `.scenario.md` files name the tool that regenerates them.** They now point at `kerno_endpoint_test`, in place of a tool that no longer exists.

**`kerno_guide` is replaced by bundled skills.** Its nine topics now ship as skills, among them `kerno-endpoint-test-intent`, `kerno-presenting-diffs` and `kerno-reviewing-lessons`, listed and read like any other skill.

***

### September 4, 2026

#### What's new

**Test AWS serverless applications end to end.** Kerno now tests AWS SAM and Serverless Framework services. It lists their HTTP routes, discovers SQS-triggered Lambdas as background consumers and runs them through the full test cycle, reads and seeds DynamoDB so tests can assert on the records a consumer wrote, and signs in to routes guarded by an AWS Cognito user pool.

**Each repository gets its own Kerno organization.** The first time you use Kerno in a repository it asks which organization to work in, or lets you switch Kerno off there entirely. The active organization shows in the status bar, so work in a side project is no longer attributed to your work account without you knowing.

**Test a whole set of endpoints in one call.** A new batch command dispatches many endpoint tests at once and bounds how many run in parallel, with one more call to wait for the whole batch to finish. Every finished test now returns a machine-readable verdict, whether the application, the scenario, or the environment was at fault, the per-scenario results, and how many passed.

**Sign in from CI with an API key.** You can now authenticate headless with `kerno login --api-key`, or by setting `KERNO_API_KEY`, so CI runners and scripts sign in without a browser. Keys are created in the portal under Settings, one per user per organization, with an optional expiry.

**Run your baseline tests in CI.** A new GitHub Action replays the baseline tests you have committed against your running application on every pull request, and reports them as a check with a job summary and JUnit output. It needs no Kerno account, no API key and no agent: it pulls one public image and runs the tests already in your repository. Baselines are still created on your machine, where you can review them before committing. The API key above is for running the Kerno CLI headless; the Action itself needs no credentials. See [How to Run your Baseline Tests in CI](/docs/guides/run-tests-in-ci).

#### Improvements & Bug fixes

**A dead application no longer reports a passing run.** A fault raised before planning is now recorded as a terminal failure instead of leaving the previous run's success standing, and infrastructure faults are named rather than arriving as unclassified.

**An endpoint that does not exist is rejected up front.** Asking to test a path the application does not serve now fails immediately, with the nearest matches, instead of being accepted and failing minutes later.

**NestJS routes no longer list twice.** Applications that combine a global route prefix with per-controller prefixes were listing every route twice under doubled paths, and on some apps the real paths dropped out of the analysed set entirely so nothing there could be tested. Both are fixed.

**`kerno login` works on Windows.** The authorize URL was being handed to `cmd.exe`, which read its `&` as a command separator and dropped every parameter after the first, so login failed with a server error. It now opens the full URL, and the flashing console window is gone.

**`kerno status` and `kerno login` no longer crash in CI.** Running either where stdin is not a terminal used to crash with a raw-mode error, which was every CI runner. Both now run cleanly.

**Retired the legacy MCP tools.** The old environment and baseline tool surface, eleven tools including `kerno_start_environment`, `kerno_plan_baseline`, and the `kerno_compose_*` family, has been removed now that the unified endpoint-test flow covers them.

***

### August 21, 2026

#### What's new

**Test background jobs and queue consumers end to end.** Until now Kerno tested the parts of your app that answer a request. It can now test the parts that run in the background. Kerno discovers your Celery tasks and message-queue consumers, dispatches one to a real worker, waits for it to run, and checks what the worker actually did, including any message it publishes onward to a Redis stream, a Redis list, or a Celery event. This extends the message-queue support from earlier this month to the workers on the other end of the queue, so a Django plus Celery service gets the half that does the work covered, not just the HTTP layer that hands it off.

**Test endpoints that call other services.** You can now declare the HTTP services your app depends on, and Kerno reads their OpenAPI contracts to ground its tests. When your billing service calls a payment service to charge a bill, Kerno knows that downstream contract, so it writes real failure-mode tests instead of inventing a request that 404s or marking the scenario blocked. Declare downstream services from `kerno_save_config` or the IDE, and Kerno resolves each service's contract on demand.

**Environment healthchecks, authored and monitored for you.** Calibration now takes an inventory of your infrastructure and has an agent write a healthcheck tailored to your environment. Kerno then runs that check continuously and surfaces environment health with clear cause codes, so a database that is down or a service that never came up is visible up front and throughout a run, rather than showing up as a confusing test failure halfway through.

#### Improvements & Bug fixes

**Run as many endpoint tests at once as you ask for.** A hidden cap silently queued every endpoint test past the fourth for a given app, so asking for twenty got you the throughput of four with no signal. The cap is gone. How many to run at once is your call.

**Faster, steadier model throughput on large runs.** The way Kerno paces its model requests was reworked so planning and implementation draw from one shared budget instead of quietly competing, long runs stop stalling in an invisible queue, and a new stall timeout ends only calls that have genuinely stopped making progress while long, healthy responses run to completion.

**Concurrent endpoint tests no longer cross wires.** When many tests ran against one app at the same time, a scenario could be recorded as passed against a response to a request it never made, and the environment healthcheck could report a confident failure about an environment nothing had examined. Both are fixed.

**Your agent can tell a paused gate from a hung job.** When an agent drives Kerno over MCP, the async job tools now distinguish a run waiting on your approval from a run that is stuck, with a new `job_lifecycle` guide and clearer errors when a required argument is missing.

**MCP servers list their tools in the IDE.** An MCP-serving app now shows its actual runtime tools in the endpoints panel instead of only an incidental health route.

**Validating changed endpoints has its own command again.** `kerno_validate` now takes an explicit list of endpoints to re-test, and `kerno_endpoint_test` focuses on a single endpoint. One tool for one endpoint, one tool for orchestrating a set.

**Monorepo endpoint listing restored.** Selecting a file in a monorepo project now lists its endpoints, the gutter markers and the inline add-and-run action are back on every endpoint file, and a file that fails to open now reports the error instead of doing nothing.

**More memory headroom on large repositories.** The agent's memory ceiling was raised from 1 GB to 2 GB, so indexing and test runs have more room on big codebases.

**A failed environment login is reported honestly.** When Kerno cannot log in to your running app, it now classifies that as an environment fault rather than scoring it as a test failure.

**Sharper change detection on chained routes.** Editing a shared route no longer mis-attributes the change across chained routes or duplicates endpoints when working out what to re-test.

***

### August 7, 2026

#### What's new

**One-time calibration onboards Kerno to your repo.** This is the biggest change in how Kerno starts on a new codebase. The first time you point Kerno at a repo, it runs a single calibration pass that captures the facts every endpoint test shares, how to authenticate as a real user, a working test credential, how to seed data the app's own way, what makes a write fail for reasons unrelated to the endpoint, and, for multi-tenant apps, a second identity for isolation tests. Kerno checks each fact against your running app, confirms it with you one at a time, and saves the confirmed set for the repo. From then on every endpoint test reads that context automatically, so your first real runs land clean rather than failing partway through on a bad credential or a hidden write blocker. Calibration runs once per repo and sets the foundation every later test builds on.

**Test endpoints that run on message queues.** Kerno can now test endpoints that publish to or consume from a message broker, with support for Kafka, RabbitMQ, Amazon SQS, and Azure Service Bus. A test can trigger your endpoint, wait for the message it produces or consumes, and check that the right message landed on the queue.

**Test your MCP servers like any other endpoint.** If your service exposes tools over MCP, Kerno now tests those tools through the same plan, implement, and validate flow it uses for HTTP endpoints, and returns the same pass, fail, and blocked verdicts. A running MCP service can serve as the system under test.

**Deeper Django and Django REST Framework support.** Kerno now recognises Django and DRF conventions when it discovers and tests endpoints, so it maps your routes and their expected inputs and outputs more accurately. Large Python and Django codebases index reliably.

#### Improvements & Bug fixes

**Session-cookie authentication.** Kerno can now test endpoints protected by session cookies, signing in and carrying a valid session through the run, alongside the token-based auth it already supports.

**Database schema from your Mongoose models.** Kerno reads your MongoDB data model straight from your Mongoose schemas, so Mongoose-backed apps generate database scenarios with an accurate picture of your collections.

**Automatic repository context.** During analysis Kerno now builds a picture of your repository, its structure, conventions, and dependencies, and uses it when planning and generating tests, so tests reflect how your project is actually built.

**Updating a suite re-baselines only what drifted.** When you run an update on an implemented suite, Kerno now re-baselines the scenarios that actually changed, adds coverage for genuinely new logic, and keeps everything else as-is, rather than rewriting the whole suite.

**Faster endpoint discovery and syncs.** Listing endpoints for the first time and re-syncing after code changes are both noticeably faster, and the endpoint list now fills in progressively as Kerno finds them.

**Steadier code indexing.** Kerno now waits until indexing is genuinely ready before reporting healthy, and large repositories index reliably without running out of resources.

**Sharper endpoint code context.** Kerno now reads the right code for each route when it generates tests, including endpoints it could not detect automatically.

**Tighter context handling on long runs.** Long runs now manage their context more accurately and stay within budget.

**The IDE keeps its endpoint list fresh.** The extension now refreshes the endpoint list periodically so it reflects the latest analysis.

***

### July 24, 2026

#### What's new

**Security testing that catches what SAST and PR review miss.** SAST reads your code and PR review reads your diffs, but neither reliably catches runtime failures such as broken authorization, missing access checks, or input handling that looks safe in source and breaks at the endpoint. Kerno now generates OWASP-aligned security scenarios, executes them against your live endpoint, and returns evidence to your AI coding agent, so the fix loop stays in the same session instead of triaging vulnerabilities after merge. Security scenarios ship next to your functional tests: one report, one verdict pipeline, no separate DAST tool to run or reconcile. Update your agent to extend testing to security.

**Effort modes let you trade off speed and depth per run.** A new per-run dial sets how fast a run returns against how much it covers, instead of living with a single preset, and it works the same in black-box (HTTP-only) or white-box (with datastore read and write) mode. Low Effort passes each scenario once for a fast, compact signal. Medium Effort passes each scenario twice in a row to prove it is genuinely repeatable. High Effort, the default, adds the adversarial agent's review on top for the strongest assertions and the deepest coverage.

**Your agent gets told the next step, not the menu.** Environment responses now carry a `next_action` field pointing at the exact call that should come next, and the application list carries a `recommended_next_action`, plus enumerated gate and terminal statuses, so your agent never invents its own idea of "done" and can drop into the loop mid-flow and still find its footing.

**Testing preferences save to your repo, not to Kerno.** When Kerno defaults an effort or label because your agent did not specify one, the response nudges the agent to persist the answer to your own CLAUDE.md, .cursor/rules, or AGENTS.md. Preferences stay diffable, auditable, and travel with your code, and Kerno stays stateless about your habits.

#### Improvements & Bug fixes

**A stub scenario now reports as not implemented, not passed.** Scenarios with no assert used to score as passed, which inflated pass counts and hid endpoints that had no real coverage. They now report as `not_implemented` end to end, across the agent, the OpenAPI schema, and the run-report UI.

**Multi-assert scenarios no longer drop asserts.** Scenarios with more than one assert block were only running the last one, so a real regression could still score a pass. Every block runs now.

**Baseline mismatches no longer stop at the first difference.** Every assertion in a scenario now runs even after one fails, so a single run reports every difference it found, each with its own clue and diff, instead of hiding the rest behind the first one. Closed-world assertions become the destination for scenarios that need a hard contract.

**Unfulfillable preconditions now fast-fail.** NOT NULL and constraint seed failures are classified as environment faults and escalated after one or two attempts. Previously the loop retried the same failing insert 18 to 24 times over roughly 19 minutes.

**Black-box mode no longer needs database access to pass the readiness probe.** Set a system-under-test URL, skip the datastore block, and tests run. The precondition gate no longer assumes a `DATABASE_URL` is available in HTTP-only runs.

**Drizzle ORM projects index correctly.** The schema-derivation guardrail now recognises `drizzle/` migrations alongside a TypeScript schema file, so it no longer blocks database scenarios on Drizzle apps.

**Login storms no longer trip your rate limiter.** Parallel scenarios now share a client pool instead of each re-authenticating against your login endpoint, so auth rate limiters stay quiet and scenarios stop failing with a 429 for the wrong reason.

**Preconditions no longer leave data behind.** They now use your app's own creation path, its ORM or HTTP endpoints, and honour `test_generation_context` for pre-seeded fixtures, so there are no more 409 conflicts on the second run of an endpoint.

**"Potential bug" replaces "BUG DETECTED".** Findings now carry better evidence and can be dismissed with a persisted reason, which cuts false-positive noise.

**codegraph init works on Node 20.** A first-run failure that blocked all codegraph MCP tools has been fixed.

**Headless workspace switching.** Non-interactive `kerno init -w` now takes a `--force-switch` flag so agents can switch workspace without a TTY prompt.

**Endpoint analysis no longer crashes on non-integer line numbers.** Progress is no longer lost mid-analysis when endpoint metadata contains headers or OpenAPI examples.

***

### July 15, 2026

#### What's new

**Agent memory: Kerno learns as it goes.** Kerno now keeps a per-endpoint, git-backed memory of what it learned on previous runs and shares it across runs and teammates. It adapts and improves over time instead of starting from scratch on every run, and because the memory lives in your repository the whole team benefits from it.

#### Improvements & Bug fixes

**Honest results: blocked instead of a false pass.** When a scenario cannot actually be run, Kerno now reports it as blocked rather than silently rewriting it and scoring it as a pass. Every run ends with one clear pass, fail, or blocked outcome.

**Smarter, self-correcting planning.** The planner now inspects your code and running environment and corrects itself instead of guessing, and you can track, re-run, or escalate individual scenarios instead of regenerating everything.

**Plugin install works in headless agent shells.** Installing the Kerno plugin from a non-interactive agent shell no longer crashes, so the two-step install works everywhere your agent runs.

**Reworked billing and usage reporting.** The billing screen has been rebuilt, and the portal now shows issues caught and links each run's report from the usage table.

**Onboarding and sign-in fixes.** New users reliably land in onboarding, verification emails send at registration rather than only after a resend, and dev sign-in no longer fails with "Could not load your organizations."

**Cancelling a test is clean.** Cancelling no longer breaks that endpoint or leaves a dangling prompt, free-text questions from the agent are answerable from your agent, and a cold first run gives a clear message instead of a misleading "run generate first."

**No login storms tripping rate limits.** Multiple scenarios no longer hammer your app's login and trip rate limiters, which had been causing false failures.

**IDE workflow polish.** Reports open automatically on completion with an explain-the-diff view, plan review is a self-contained page, deleted tests disappear from the list, and there is less flicker with clearer progress spinners.

***

### July 10, 2026

#### What's new

**Kerno plugin for Claude Code, Cursor, and Codex.** You can now install Kerno as a plugin from the marketplace in two steps. Your coding agent drives Kerno end to end with no manual MCP setup.

*Retired in September 2026. Kerno sets up through the CLI and MCP server now, with no plugin to install. See the* [*Quickstart*](/docs/getting-started/quickstart)*.*

**Custom rules.** You can now define your own rules that guide how Kerno plans and generates endpoint tests. Kerno follows your team's conventions and priorities instead of a one-size-fits-all default.

**Authenticated endpoint testing.** Kerno can now test endpoints that sit behind authentication. It mints valid tokens, grounds them against real baselines, and uses your validated credentials during test generation.

**Database schema straight from your source.** Kerno now reads your database schema directly from your code, so it plans and generates tests with an accurate picture of your data model without any manual setup.

**S3-compatible object storage support.** Services that read from or write to S3-compatible object storage, such as MinIO, can now be tested against a real running instance that Kerno provisions and tears down for you.

#### Improvements & Bug fixes

**One command to start: `kerno init`.** The separate `kerno start` and `kerno mcp` commands are now a single `kerno init`, and it runs cleanly in headless and agent shells that previously crashed.

**One browser window to sign in.** Logging in from the IDE or CLI no longer opens the browser twice. Authentication and organization selection are consolidated into a single window.

**Steadier long runs.** The model connection is now kept alive so long calls no longer stall silently, and large-repo indexing that used to hang now completes.

**Faster re-analysis after edits.** Kerno now re-analyzes only what changed after a code edit, following the call graph, instead of recomputing whole modules. Repeat runs on large repositories are noticeably faster.

**Credentials handled correctly.** The agent now uses the credentials you provide instead of faking a pass or making destructive database writes, and it asks for them up front in black-box mode.

**MCP sessions survive agent restarts.** Restarting your agent no longer silently drops the Kerno MCP session, and duplicate log noise on connect has been removed.

***

### July 3, 2026

#### What's new

**Signature-verified webhook endpoints.** Endpoints protected by request signatures, such as Stripe and GitHub style webhooks, are now testable. Kerno signs requests correctly so these routes can be validated like any other.

**Preconditions: Kerno sets up required state before testing.** Kerno now treats the setup a scenario needs, for example creating a record before fetching it, as a first-class step. More endpoints can be tested reliably without manual staging.

#### Improvements & Bug fixes

**More accurate change detection.** A wave of fixes means changed string fields, value changes inside arrays of objects, and differences in nested responses are now caught reliably. Per-run values such as ids and timestamps no longer show as false failures, binary and non-JSON responses no longer flag on a no-op re-run, and the diff compares the real response instead of a cleanup response.

**"What changed?" re-tests the right endpoints.** Editing a shared service or helper file now re-tests the endpoints that actually depend on it, following the call graph, instead of the wrong endpoint or none.

**Control at the scenario level.** You can now see per-scenario progress, re-run a single scenario, and escalate one for review without regenerating everything for that endpoint.

**No accidental loss of passing tests.** Re-planning no longer deletes existing passing scenario files before you approve the new plan.

***

### June 26, 2026

#### What's new

**Point local endpoint tests at a real database.** You can now configure a database for local endpoint tests, so scenarios run against real data instead of being limited to black-box calls.

**Configure your dependencies from the IDE.** A new configuration screen lets you set up external dependencies, such as Postgres, Redis, and MongoDB, directly from the extension in VS Code and JetBrains.

#### Improvements & Bug fixes

**Reliable analysis on large, real-world codebases.** Very large monorepos previously stalled for 20 minutes or more, or ran out of context during setup. Endpoint discovery and analysis now stay within limits and complete reliably on million-line repositories instead of dead-ending.

**Broader dependency coverage.** Kerno now handles the full set of supported external dependencies consistently during environment setup.

**Deeper, module-aware security analysis.** Endpoint security analysis now considers the surrounding module, not just the endpoint in isolation, so it catches issues that only show up in context.

**Recap labels match the real results.** The pass and fail summary at the end of a run could previously disagree with the per-scenario outcomes. The recap now reflects what actually happened.

**Security findings report real values.** Findings now reference the actual environment variable names Kerno found instead of fabricated placeholders, so they are directly actionable.

**Automatic context summarization on large runs.** When a run's context grows past a safe threshold, Kerno now summarizes earlier steps automatically instead of overflowing and failing.

***

### June 19, 2026

#### What's new

**Choose how your environment is set up.** A new setup mode selector lets you tell Kerno whether to run your app locally or point at a remote running instance, so Kerno fits how your project already runs. Available in VS Code and JetBrains.

#### Improvements & Bug fixes

**See whether your app is reachable.** The extension now shows whether your system under test is running and suggests starting your local app when it cannot be reached, so setup issues are obvious up front.

***

### June 12, 2026

#### What's new

<figure><img src="https://3940381227-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FuNjVcN7yphTDowu96vxO%2Fuploads%2FGvXsFMGDu6y1deQZYqrW%2Funnamed.gif?alt=media&amp;token=d4a16eaf-5387-4cee-a22f-ebf21570a477" alt=""><figcaption></figcaption></figure>

**Rebuilt IDE experience.** The extension is now a visual aide that reflects what the agent is doing rather than a control surface you drive Kerno from. Environment composition, baseline tests, and endpoint validation are triggered via MCP and appear in the extension as they happen instead of waiting on a silent run.

**Agent and engineer-in-the-loop SUT builds.** When a system-under-test build hits something Kerno cannot resolve alone, it reports what it tried and asks a targeted question: where a definition lives, which config it needs. You or your agent answer inline and the build continues instead of dead-ending. This is a one-time onboarding action per repo. Once the `.kerno` folder is committed, subsequent builds reuse the existing compose files instead of starting from zero.

<figure><img src="https://3940381227-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FuNjVcN7yphTDowu96vxO%2Fuploads%2FsD8yZ2INysFoI0IINTLb%2Funnamed.png?alt=media&amp;token=f0548aff-0840-44c5-a548-ca835faeac9e" alt=""><figcaption></figcaption></figure>

**Guided setup for new projects.** Projects that are not yet configured get a walkthrough so the first build starts from a known-good state instead of dead-ending on missing setup.

**Clear Cache from the IDE or CLI.** Reset a workspace from Settings → Danger Zone, or run `kerno reset` from the CLI. Clears both the repo copy and the temporary workspace copy so a stale scenario cannot leave the agent reporting an outdated run.

**Bedrock model fallback.** An automatic fallback path for model calls reduces run failures when a primary provider is unavailable.

#### Improvements & Bug fixes

**Copy-prompt flow replaces auto-regeneration.** Kerno no longer regenerates plans automatically. You copy the prompt and decide when to send it to your agent.

<figure><img src="https://3940381227-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FuNjVcN7yphTDowu96vxO%2Fuploads%2FCm7z0togyD1g0CtWZxjh%2Funnamed%20(1).png?alt=media&amp;token=89e27ff3-d4dc-4305-aee0-d35c364f252d" alt=""><figcaption></figcaption></figure>

**More reliable endpoint discovery.** Fixed a code-graph hashing inconsistency that caused endpoints to be re-discovered between runs.

**Sandbox TLS fix.** Resolved a TLS failure in the TypeScript sandbox that was reducing test run reliability.

**Prompt templates and markdown rendering.** Reusable prompt templates and cleaner markdown rendering in the extension.

***

### June 4, 2026

#### What's new

**Live MCP activity in the extension.** When Claude Code, Cursor, Windsurf, or any other MCP host drives Kerno, the extension now reflects what is happening in real time. Test runs, scenario plan and implement, environment spin-ups, baselines, and validation triggered by your agent all show up in the extension within a few seconds.

**Live status in the Tests view.** Per-endpoint running indicators now appear when your agent is planning, implementing, or validating. Coverage and the report panel refresh automatically when the run completes so you always see the latest state without having to reload.

**Notifications for agent-triggered runs.** Completion toasts now fire for runs started by your agent outside the IDE, not just for runs you triggered manually from the extension. You know when an agent's run finishes without having to watch the terminal.

**MCP connectivity indicator.** The Environment panel now shows a live MCP port and status dot. You can see at a glance whether an agent is connected and whether it is currently doing something.

#### Improvements & Bug fixes

**Approve-and-update no longer leaves stale diffs.** Approving a diff was clearing the decision but leaving the report showing "2 diffs detected." The report now clears correctly after approval.

**"Test complete" no longer fires during planning.** The completion toast was firing as soon as planning finished, before implementation had run. The toast now waits for the full run to complete.

**Spinner on the Implement Plan button.** Long implementations previously showed a greyed-out button with no feedback. There is now a clear in-progress spinner so you know the operation is running.

**Stale environment after plan feedback.** Providing feedback on the compose plan and regenerating it was not triggering a re-run of orchestration. kerno\_start\_environment now re-runs orchestration after a plan change instead of reusing a cached healthy result.

**Validation no longer hangs without a Docker healthcheck.** Environments with no healthcheck configured were stalling at the ready-for-validation step indefinitely. These environments now reach ready state correctly.

***

### May 21, 2026

#### What's new

**Azure Blob Storage and Azure Data Tables via Azurite.** Kerno now automatically orchestrates and tests against Azure Storage services. When Kerno detects an Azure Storage connection string in your app, it provisions the Azurite emulator, injects a working local endpoint, and wires it into your test environment. State is wiped between runs and the container is health-checked before any scenario starts. Both Blob and Data Tables are supported.

**ts-sandbox upgraded to Node 24 LTS.** The TypeScript test sandbox now runs on Node 24. This brings native fetch, the updated V8 engine, and keeps the runtime ahead of Node 20's maintenance window. No changes required on your side.

**Scenario agent now inspects code and environment before retrying.** Two new analysis tools give the implementation agent direct visibility into your codebase and running environment. Instead of burning retries on the wrong assumption, the agent now exits early with a useful diagnostic when it hits a problem it cannot resolve.

**Review the environment plan before building.** Kerno now renders the full Docker Compose plan as markdown before spinning anything up. Approve with one click, give free-text feedback to regenerate, or cancel. The plan waits for your sign-off before any container is started.

**Add or remove services from the plan.** Services added or removed through the environment selector stay in sync with the plan view. You can also adjust services directly from the plan screen before approving the build.

#### Improvements & Bug fixes

**npm install fix on Linux.** Installing @kerno/cli was crashing on Linux with an EBADPLATFORM error caused by a macOS-only fsevents dependency being pulled in unconditionally. The dependency is now gated correctly and Linux installs complete cleanly.

**Extension login on Linux.** The extension was failing on startup with a `spawn secret-tool ENOENT` error when the OS keychain helper was not available. This is now handled gracefully and Linux users can log in without a workaround.

**Agent version visible in the IDE.** The current agent version now displays in the extension panel. When a newer version is available, an in-place update prompt appears so you can upgrade without leaving your editor.

**Kerno asks when setup gets stuck.** The environment orchestrator previously stopped with a generic infra error when it encountered an ambiguous situation (wrong DB driver, missing env var, unsupported config). It now escalates to a user-feedback prompt so you can provide the missing information and continue.

**Code changes picked up correctly on re-run.** Modifying code and re-running Kerno was not producing updated results after the initial run. Changes in the working tree are now detected correctly on every run.

**Failed run state no longer sticks between sessions.** Bad state from a failed build was being written back to .kerno/ and causing a persistent BuildSUT retry loop on the next session. Bad state is now discarded instead of persisted.

**Closing the IDE now cancels test generation.** The backend now auto-cancels the agent run on client disconnect. Previously, closing the IDE left orphaned runs spending tokens in the background.

**Scenario generator no longer produces uncompilable TypeScript.** Multi-line string literals in generated .scenario.ts files were causing parse errors. The generator now handles these correctly and all generated files compile on the first attempt.

***

### May 7, 2026

#### What's new

**Instantaneous SUT rebuild on code change.** When you modify code, the SUT environment now rebuilds immediately. Previously there was a noticeable delay between saving a change and the environment reflecting it. The synchronisation is now fast enough that you can validate immediately after editing.

**Implementation agent can now diagnose environment issues directly.** The agent now has shell access inside the sandbox, which lets it inspect what is actually running rather than inferring state from error messages. When something goes wrong during implementation, the agent can check container state, logs, and process health before deciding how to proceed.

#### Improvements & Bug fixes

**Env build failures now surface a prompt.** When Docker Compose fails during environment setup, Kerno previously stopped with an infra error and gave no path forward. It now asks for more information and offers remediation options so you can correct the issue without starting from scratch.

**Kerno stop now cleans up completely.** The agent process and all running Docker containers now fully tear down when you stop Kerno. Previously, some containers were left running in the background after a stop.

***

### April 16, 2026

#### What's new

**Large repo support.** Kerno now handles repositories with over one million lines of code reliably. Index timeouts that previously caused large codebases to fail during setup have been resolved, and the indexer now completes successfully on projects of any size.

**Workspace analysis caching.** Analysis results are now cached correctly across runs. If your codebase has not changed since the last analysis, subsequent runs skip the redundant work entirely and proceed directly to the operation you requested.

#### Improvements & Bug fixes

**Auth tokens now refresh through long runs.** Tokens were expiring during long implementation sessions, causing authentication failures mid-run. Tokens now refresh silently in the background so you stay signed in from start to finish regardless of how long a run takes.

**kerno stop now tears down everything.** The stop command previously left the agent process or Docker containers running in some cases. The agent process and all containers now fully shut down on stop.

**File writes succeed on paths containing dots.** File paths with dots in directory names were being percent-encoded incorrectly, causing writes to fail with a path not found error. Paths are now handled correctly and writes land on the first attempt.

**Repeat runs skip redundant work.** Workspace analysis results were not being reused between runs even when nothing had changed. The cache now works correctly, so repeat validation runs are significantly faster.

***

### April 9, 2026

#### What's new

**Codex is now a supported MCP agent.** Add Kerno to your AGENTS.md and Codex can index your codebase, compose the environment, and run full validation. Claude Code and Cursor are also supported. All three agents use the same MCP interface.

**Scenarios tab in the IDE.** A dedicated tab now shows all generated scenarios for the selected endpoint. You can inspect each scenario individually and re-run specific ones without triggering a full regeneration of the entire endpoint.

**Agents can no longer skip the environment build plan.** The MCP tool schema now enforces that kerno\_start\_environment cannot be called until the compose plan has been shown and explicitly approved. This prevents agents from bypassing the review step, which catches edge cases that Kerno cannot detect automatically such as a shared database that already exists in the repo.

#### Improvements & Bug fixes

**Environment status now reflects SUT readiness accurately.** kerno\_start\_environment was previously returning "up" while the application container was still initialising or had failed silently. The status response now distinguishes between dependencies running and the full application being ready for validation. You will no longer see false positives on environment health.

**Large repo indexing no longer times out.** Repositories above a certain size were hitting index timeouts that caused setup to fail. The indexer now handles large codebases without failing on first run.

***

### April 2, 2026

#### What's new

**ClickHouse support.** Kerno can now spin up ClickHouse as part of the test environment. If your service reads from or writes to ClickHouse, scenarios run against a real running instance rather than a stub. Kerno provisions the container, waits for it to be healthy, and tears it down cleanly after the run.

**Auth tokens stored in OS keychain.** VS Code, JetBrains, and the CLI now store authentication tokens in your operating system keychain and refresh them silently before they expire. You stay logged in across sessions without having to re-authenticate after restarts.

#### Improvements & Bug fixes

**Sign out now works correctly.** Clicking sign out was redirecting back to the dashboard instead of signing the user out. The sign out flow now completes correctly and clears the stored session.

**Re-indexing during Cursor edits no longer breaks the environment.** Editing files in Cursor while Kerno was re-indexing could leave the environment in a broken state. The indexing and environment processes are now properly isolated.

***

### March 19, 2026

#### What's new

**Python PEP 735 dependency groups supported.** Python projects using PEP 735 dependency groups in their pyproject.toml were failing to index correctly on the first run. Kerno now parses these correctly and more Python projects index successfully without any manual intervention.

**Mocking is more reliable.** Synthetic route registration was failing silently in certain configurations, causing a retry loop that did not surface a useful error. The registration flow has been hardened and failures now surface an actionable error message on the first attempt.

#### Improvements & Bug fixes

**Coverage metric fix.** An endpoint was being marked as covered as soon as a scenario plan was created, even before any scenarios had run. Coverage now only increments when scenarios have actually executed successfully.

**Faster environment startup on repeat runs.** Kerno previously rebuilt the Docker environment from scratch on every run for a repository it had already set up. Kerno now reuses what it built and only rebuilds what changed, significantly reducing startup time on subsequent runs.

**Python indexing improvements.** Several edge cases in Python project indexing were causing silent failures. These have been resolved and the indexer now handles a broader range of Python project structures correctly.

***

### March 12, 2026

#### What's new

**Schema and documentation generation.** Kerno now generates accurate API schemas and documentation directly from your codebase. These are generated from the actual code rather than hand-written annotations, so they stay in sync as your code changes. Your agent can use these as a reliable source of truth when reasoning about your API.

**Log in once, stay logged in.** VS Code, JetBrains, and the CLI now store authentication tokens in your OS keychain and refresh them silently before they expire. You no longer need to re-authenticate after restarting your editor or machine.

#### Improvements & Bug fixes

**Stopping Kerno now fully stops everything.** Running kerno stop previously left agent processes or Docker containers running in some configurations. The stop command now shuts down the agent process, all Docker containers, and the environment cleanly every time.

**Mocking failures now surface errors immediately.** Persistent failures in synthetic route registration were causing silent retry loops. These now fail fast with a clear error message rather than retrying indefinitely.

***

### March 5, 2026

#### What's new

**Kerno MCP.** Kerno is now headless and works natively with your coding agent. Kerno MCP exposes Kerno's full capabilities directly inside Claude Code, Cursor, and any MCP-compatible agent. Your agent can search your codebase, start environments, and run full end-to-end validation without you touching the IDE.

**kerno\_validate.** Your agent runs the full integration test suite after making a code change. No context switch and no manual trigger. The feedback loop closes in the same agent session, so your agent can write code, validate it, and fix issues without leaving the conversation.

**kerno\_env.** Your agent spins up databases, queues, and service dependencies on demand. The real services, running locally, ready to test against. Your agent no longer needs you to manage the environment manually.

**Codebase search tools.** MCP exposes Kerno's semantic code index to your agent. Your agent can search across endpoints, schemas, and dependencies and get precise answers in a single tool call without burning tokens on broad file reads.

#### Improvements & Bug fixes

**Faster environment startup on repeat runs.** Kerno now reuses previously built containers and Docker layers instead of rebuilding from scratch. Environments that previously took several minutes to start on repeat runs now start in seconds.

***

### February 26, 2026

#### What's new

**Kerno CLI (early access).** The Kerno CLI is now available for early adopters. Install globally via npm and run Kerno from any terminal on macOS, Linux, or Windows without needing the IDE extension.

```bash
npm install -g @kerno/cli
```

**Proper exit codes for CI/CD.** Failed CLI commands now return meaningful non-zero exit codes. Kerno integrates cleanly into automated pipelines and CI systems that rely on exit codes to detect failures.

#### Improvements & Bug fixes

**Sign out actually signs you out.** Clicking sign out in the IDE extension was redirecting back to the dashboard rather than ending the session. The sign out flow now completes correctly and clears all stored credentials.

**Agent processes clean up on exit.** Closing the terminal or interrupting the CLI process now shuts down the agent and any associated Docker containers cleanly, rather than leaving orphaned processes running in the background.

***

### February 19, 2026

#### What's new

**Review the plan before anything runs.** Kerno now shows you the full scenario plan before starting implementation. You can approve it as-is, reject individual scenarios, or give feedback to adjust scope and direction. Nothing is generated until you sign off.

**Stop a run mid-implementation.** You can now stop a long-running implementation at any point. Make your code change, then start again from a clean state. There is no broken state left behind and no error on retry.

**Update scenarios after the fact.** If you change your mind about a scenario after implementation, you can edit and re-run it individually without triggering a full regeneration of all scenarios for that endpoint.

#### Improvements & Bug fixes

**Run state no longer persists across restarts.** Failed runs were leaving partial state in .kerno/ that caused errors on the next session. State from failed runs is now discarded on exit so each session starts clean.

**Scenario generator no longer produces uncompilable TypeScript.** Multi-line strings in generated scenario files were being formatted incorrectly, causing TypeScript parse errors. Generated files now compile correctly on the first attempt.

***

### February 12, 2026

#### What's new

**Multi-org support.** You can now belong to and switch between multiple organisations in Kerno. Teams that work across more than one organisation no longer need separate accounts.

**AI thinking logs.** During environment setup and baseline capture, you can now see the agent's reasoning as it works. Thinking logs are displayed inline so you can follow what the agent is doing and catch issues early rather than waiting for the operation to complete.

**Enable and disable Kerno from the IDE.** You can now toggle Kerno on and off from the extension panel without uninstalling. Useful for pausing Kerno on a project temporarily without losing your configuration.

#### Improvements & Bug fixes

**Median latency reduced to 50 seconds.** Index time is down from a median of 110 seconds to 50 seconds, measured across 33 runs on TypeScript and Python projects ranging from small to large. Minimum observed is 20 seconds.

**Clear Kerno logs from the IDE.** The extension panel now includes a clear logs action so you can reset the output view without restarting the extension.

**Onboarding walkthrough improved.** The initial setup flow now guides you through each step with clearer instructions and better error messages when something is missing.

***

### January 29, 2026

#### What's new

**In-IDE behaviour diffs.** Kerno now shows behaviour diffs directly in the IDE panel. When a test run detects a change, you can inspect the expected versus actual response inline without leaving your editor.

**Enable and disable Kerno per project.** You can now turn Kerno on or off for a specific project from the IDE panel. Kerno remembers the setting per project so you do not need to reconfigure on each restart.

#### Improvements & Bug fixes

**Dozens of small improvements.** This release contains a broad set of tweaks that reduce cognitive overload during test generation and review: clearer status messages, more consistent UI state, and better handling of edge cases during environment setup.

**Clear Kerno logs action.** The extension panel now has a clear logs button so you can reset the output view without restarting.

***

### January 22, 2026

#### What's new

**New portal.** A new web portal is now available at portal.kerno.io. It shows endpoint coverage across your app and tracks P0 issues detected by Kerno. Teams can use the portal to monitor coverage progress and review outstanding issues without opening the IDE.

**VS Code native UI.** The IDE extension has been rebuilt using VS Code components. The experience now feels native to the editor rather than a custom-styled panel embedded inside it.

#### Improvements & Bug fixes

**Polyglot support improvements.** TypeScript and Python projects across a wider range of configurations now index and run correctly. Several edge cases that caused silent failures on project setup have been resolved.

**General stability.** A broad set of small fixes across environment setup, test generation, and result reporting. The overall number of unexpected errors during a session is significantly reduced.

***

### January 15, 2026

#### What's new

**Just-in-Time containerisation.** Kerno now generates Dockerfiles and Docker Compose manifests on the fly for services that do not have them. This happens automatically as part of environment setup with no manual configuration required. The generated files update automatically as your service and API logic changes.

**Automatic dependency mapping.** Kerno maps your internal and external dependencies, migration context, and business logic during setup. This gives the test agent the context it needs to generate accurate scenarios without you having to explain the architecture.

#### Improvements & Bug fixes

**Agent CPU and memory requirements down by 50%.** The Kerno agent is now significantly more lightweight. Running Kerno alongside your editor and application no longer competes noticeably for system resources.

**Average test time reduced to \~2 minutes.** From spinning up the backend service and seeding dependencies to writing and executing tests, the full cycle now completes in around two minutes on most projects.

**Setup time \~3 minutes.** Initial indexation now completes in around three minutes for small to medium projects.

***

### January 8, 2026

#### What's new

**Seven infrastructure services supported out of the box.** Kerno now supports Postgres, MySQL, MariaDB, MongoDB, Redis, Kafka, and RabbitMQ as first-class test environment services. Kerno handles data and context propagation across all dependencies automatically. No migration files to manage, no flaky Docker Compose configurations, and no wrong database version surprises.

**IDE support across five editors.** Kerno now works with Antigravity, Cline, Windsurf, Cursor, and VS Code. Install the extension for your editor and Kerno connects automatically when you open a supported project.

#### Improvements & Bug fixes

**Initial release stability.** A broad set of fixes across environment provisioning, dependency detection, and test execution landed alongside the initial release. Kerno handles the most common project configurations without requiring manual intervention on setup.


# How Kerno Works

Understand how Kerno works, from indexing your codebase to validating code changes.

Kerno is runtime code and security review. It builds a deep understanding of your codebase and captures how your code operates. Every time you change code, it checks the change against that baseline to catch regressions, integration issues, and security gaps.

Your agent drives Kerno over MCP. It runs inside your agent's loop, so your agent validates what it just wrote, reads the results, and fixes what broke in the same session.

{% code expandable="true" %}

```mermaid
  flowchart LR
      s1["#nbsp;1. Index your codebase#nbsp;"] --> s2["2. Connect app"] --> s3["3. Generate tests"]
      s4["4. Review code changes"] --> s5["5. Update tests"] --> s6["6. Learn and improve "]
```

{% endcode %}

{% stepper %}
{% step %}

#### Index your codebase <i class="fa-code">:code:</i>

Kerno analyzes your code and builds a graph of every function, class, model, and entry point, and how they connect. The graph gives Kerno the exact blast radius of any code change, so it knows how a change affects the rest of your system. By default, Kerno shares the code index it builds with your teammates through your repository's git remote. See [Codebase indexing.](/docs/core-concepts/codebase-indexing)
{% endstep %}

{% step %}

#### Connect Kerno to your running app <i class="fa-plug">:plug:</i>

Connect Kerno to your running app, local or remote, so it can validate your code changes against your repo's real services, dependencies, and framework. Your agent then calibrates Kerno to your repo once, confirming with you how to sign in, how to seed test data, and which infrastructure Kerno may touch. See [Environment setup](/docs/core-concepts/environment-setup) and [Calibration](/docs/core-concepts/environment-setup#calibration).
{% endstep %}

{% step %}

#### Generate baseline tests <i class="fa-clipboard-check">:clipboard-check:</i>

Kerno generates tests covering your entry point's functional behavior and edge cases, plus security tests when you ask for them, then runs them against your real app to capture a baseline, a snapshot of your entry point's current behavior. It also tests your [user flows](/docs/core-concepts/user-flows), the journeys that cross several entry points, end to end. See [Scenarios and baselines](/docs/core-concepts/scenarios-and-baselines) and [Security testing](/docs/references/security-testing).
{% endstep %}

{% step %}

#### Review code changes <i class="fa-vial-circle-check">:vial-circle-check:</i>

When you change code, Kerno maps its blast radius and re-runs the affected entry points' tests against your app, comparing the results to the baseline. It shows you exactly what your change altered, so you can tell whether each difference is intended or a bug. See [Change validation](/docs/core-concepts/change-validation).
{% endstep %}

{% step %}

#### Update tests <i class="fa-sparkles">:sparkles:</i>

When a difference is intentional, Kerno updates the suite to match the new behaviour. It rewrites the tests your change affected and writes new ones for any logic you introduced, so coverage keeps pace with your code without manual upkeep.
{% endstep %}

{% step %}

#### Learn and improve <i class="fa-brain">:brain:</i>

Kerno's memory learns from every interaction with your team, so its testing gets more tailored to your codebase over time. Custom rules capture your team's best practices and apply them consistently across every entry point Kerno tests. See [Custom rules](/docs/core-concepts/custom-rules) and [Memory and learning](/docs/core-concepts/memory-and-learning).
{% endstep %}
{% endstepper %}


# Codebase Indexing

Understand how Kerno builds a deterministic index of your codebase, keeps it up to date, and why it matters for accurate validation.

Most code tools that use LLMs rely on vector embeddings, which find relevant code by similarity. That is useful but approximate, and it does not reflect how your code is actually wired together.

**Kerno indexes deterministically**. It resolves symbols and cross-references straight from your source, so every function call, import, and variable reference gets a precise location that is reproducible across runs. The same index drives entry point discovery and impact analysis.

Route discovery is not pattern matching on filenames. Kerno reads each framework's routing constructs directly, so a route registered in a nested router group or behind a decorator is found the same way an obvious one is. See [Supported Technologies](/docs/references/supported-technologies) for the frameworks covered.

When Kerno finds no entry points in an application, for example one built on a framework it does not cover, an AI agent reads the code to find them. When you ask Kerno to test a route it did not list, it searches your code for that route's handler.

By default, Kerno shares the index through your repository's git remote, `origin`. Once a commit is on `origin`, Kerno pushes the index it built for that commit there as hidden refs under `refs/index/`, and fetches the entries your teammates pushed, so each commit is indexed once for the whole team. An entry holds none of your source files. It does hold file paths, symbol names, signatures and doc comments, and anyone who can read the repository can fetch it.

Relevant code excerpts are sent to LLM providers only when Kerno needs them to reason about a specific task. See [Security & Privacy](/docs/references/security-and-privacy).

```mermaid
   flowchart TD
       src["src/"]
       src --> routes["routes/"]
       src --> services["services/"]
       src --> models["models/"]

       routes --> userRouter["userRouter.ts"]
       services --> userService["userService.ts"]
       models --> userModel["userModel.ts"]

       userRouter --> createUser["POST /users"]
       userRouter --> getUser["GET /users/:id"]
       userService --> hashPassword["hashPassword()"]
       userService --> dbQuery["db.query()"]
       userModel --> emailField["email: string"]
       userModel --> idField["id: string"]

       classDef file stroke:#3b82f6,stroke-width:2px,fill:transparent
       classDef endpoint stroke:#f59e0b,stroke-width:2px,fill:transparent
       classDef symbol stroke:#10b981,stroke-width:2px,fill:transparent

       class userRouter,userService,userModel file
       class createUser,getUser endpoint
       class hashPassword,dbQuery,emailField,idField symbol
```

Directories, files, entry points, and the symbols inside them are all nodes in one graph. Kerno traverses it in either direction: outward from a function to the entry points that reach it, or inward from an entry point to everything it touches.

### What the graph gives you

* **Deterministic blast-radius detection.** When you change code, Kerno knows which entry points, tests, and dependencies are affected. It follows the call graph outward across your whole codebase, so a change to a shared helper surfaces every entry point downstream of it.
* **No hallucinated context.** When Kerno asks an LLM to plan a test or update a baseline, it sends real code references rather than similarity matches.
* **Re-indexing on change.** When Kerno syncs after a code change, it rebuilds its index for your current code, so the blast radius reflects what is on disk.

### What gets indexed

Kerno indexes the git-tracked source files in your project, including application code, configuration, and build files. Files that are not tracked by git (dependency directories, build outputs, anything in `.gitignore`) are skipped automatically.

A new file you just created is untracked until you stage or commit it, so stage or commit it before expecting Kerno to index and validate it.

In a monorepo, Kerno detects each application in the repository and indexes them into one graph that spans the whole repo.

### Index updates

Kerno indexes your codebase on first use. After that, your AI coding agent syncs the workspace after code changes, so Kerno always runs against the latest version of your code.

You can also trigger a sync yourself. Ask your coding agent to sync the workspace

{% code expandable="true" %}

```
Sync the Kerno workspace so it picks up my latest changes.
```

{% endcode %}

Kerno takes a fresh snapshot of your code and re-runs its analysis.

{% hint style="warning" %}
A sync stops every Kerno run in progress, including test generation and change reviews. Let running work finish before you sync.
{% endhint %}


# Test Environment

Understand how Kerno connects to your running application and its dependencies to test against your real stack.

Kerno connects to your running app and the services it depends on, so it runs tests against your real stack. That gives you high-fidelity results that reflect how your app behaves in production.

```mermaid
  %%{init: {"themeVariables": {"clusterBkg": "transparent", "clusterBorder": "#cccccc"}}}%%
  flowchart LR
      subgraph yours["Your environment"]
          APP[Your running application]
          DEPS[(Databases, caches,<br/>queues, auth)]
          APP --> DEPS
      end
      SANDBOX[Kerno sandbox<br/>runs the tests]
      SANDBOX -->|HTTP| APP
      SANDBOX -.->|optional read/write access| DEPS
```

### How Kerno connects to your application

Kerno connects to your running application, and optionally to the services it depends on.

Where your application runs takes one of two shapes.

* **`local`.** The application runs on your machine, reached at the URL it listens on, such as `http://localhost:3000`.
* **`remote`.** The application runs somewhere else you can reach, such as a shared dev environment.

Before relying on that address, Kerno probes it from inside its own sandbox, so an unreachable application is caught immediately rather than surfacing later as a confusing test failure. Kerno treats the environment as ready only once that check passes, and runs entry point tests only against a ready environment.

Kerno can run tests in two modes, black box and white box. White box is the default, and it is what a run uses once you have given Kerno access to your dependencies.

#### Black box testing

In black box mode, Kerno tests your application entirely from the outside, like any external client. It talks to your application only over HTTP, and sets up and verifies everything through your own entry points, with no direct access to your dependencies.

This sets two limits. Kerno will not generate tests that rely on direct database access, and it cannot validate side effects that are not visible through your API, such as confirming what your entry point wrote to the database.

```mermaid
  flowchart LR
      S[Kerno sandbox] -->|HTTP| A[Your application]
      A --> D[(Dependencies)]
```

#### White box testing

Kerno may additionally reach those dependencies directly, for the cases your API cannot express. Two common examples are seeding a record your API has no way to create, and confirming that a stored password is not the plaintext that was submitted.

```mermaid
  flowchart LR
      S[Kerno sandbox] -->|HTTP| A[Your application]
      A --> D[(Dependencies)]
      S -->|read/write| D
      linkStyle 2 stroke:#f59e0b,stroke-width:2px
```

With that access, Kerno handles the data layer for you, reading and writing your dependencies to establish the state a test needs and to verify what your entry point actually persisted. That is the work you would otherwise write fixtures and teardown code for.

To enable it, provide connection details for the services your application depends on: its database, but also a cache like Redis, a queue like Kafka or RabbitMQ, object storage, or an auth provider like Zitadel. See [Supported Technologies](/docs/references/supported-technologies) for the full list.

Any individual run can opt out and go black box. See [Testing modes](/docs/references/testing-modes) for the dial that does it.

#### Services your app calls over HTTP

If your app calls other HTTP services, tell Kerno about them:

```
My app calls a billing service at http://localhost:4000, with its OpenAPI spec in docs/billing.yaml. Add it to Kerno's config.
```

For each service you can give its address, so tests can reach it, and its OpenAPI spec, as a URL or a path in your repository, so tests check the real request and response shapes. Kerno uses this only when it writes tests, and your app's configuration stays as it is.

### Calibration

Once your application is connected and ready, your agent calibrates Kerno to your repository. Calibration runs once per repository and records the facts every test relies on, so each test starts from them.

It works in three passes. Your agent first reads a handful of files, such as your agent instructions, README, scripts and env files. It then checks what it found against your running application with a few live requests. Finally, it walks you through each item, one at a time, and asks you to confirm or correct it. The cost stays the same however many entry points your application has.

| Item           | What calibration records                                                                                    |
| -------------- | ----------------------------------------------------------------------------------------------------------- |
| Sign-in        | How to sign in as a real user, with working test credentials                                                |
| API token      | A ready API token and its scope, or how to create one                                                       |
| Seeding        | How to create test data the way your application does, such as a seed or factory command                    |
| Write blockers | What makes a write fail regardless of the entry point, such as plan limits, rate limits or required headers |
| Tenancy        | For multi-tenant applications, a second account outside the first one's tenant, for isolation tests         |
| Infrastructure | Every database, cache, queue and store your code talks to, and whether Kerno may access each one            |

You decide what Kerno may touch. Each infrastructure component is recorded either with its connection details, or as off limits with an optional reason.

**What calibration writes.** Calibration saves its results in `.kerno/config.yaml`, which you commit:

* The confirmed facts become [custom rules](/docs/core-concepts/custom-rules) that every test reads.
* The infrastructure inventory is saved for each application.
* A `calibration` record notes the outcome, so your agent doesn't run it again.

Secret values stay out of that file. It holds only the variable name, and the value goes in `.kerno/.env`, which Kerno keeps out of git.

**The environment healthcheck.** From the infrastructure inventory, Kerno generates `.kerno/healthcheck.local.ts` (or `healthcheck.remote.ts`), with one check for your application and one for each component Kerno may access. The check runs periodically, and while it fails, Kerno reports the environment as not ready for tests. The file is yours to edit, and Kerno keeps your changes.

**Declining or running it again.** You can decline calibration at any point. Kerno records that you declined and stops asking, and tests then run without shared sign-in, seeding or tenancy setup. To calibrate later, or again after you reset your environment, ask your agent:

```
Calibrate Kerno for this repo.
```

In Claude Code, you can also type `/kerno-calibrate`. During first-time setup, the [Kerno portal](/docs/portal/overview) shows calibration progress on its setup page.

### How Kerno reads your database schema

Credentials alone are not enough to write to your database: Kerno also needs to know its shape. It reads that from your source, never by introspecting your running database. A final-state snapshot such as `schema.rb`, a `*.prisma` file or `schema.sql` is used when present, otherwise Kerno reads your migrations.

Most projects work automatically. When a schema cannot be derived, Kerno pauses and asks you for a path or for the DDL. It keeps your answer for that one entry point. Another entry point asks again, and clearing Kerno's cache for the whole workspace forgets the answer. See [Environment setup issues](/docs/troubleshooting/environment-setup-issues).

### How Kerno authenticates to your entry points

For each entry point, Kerno analyses your authentication code and records the mechanism it uses, how to present the credential, and which roles or scopes it requires. Because Kerno reads your actual authentication code, it can handle a wide range of schemes. Common ones include bearer tokens and JWTs, session tokens, API keys in a custom header, and OAuth 2.0

Where it can, Kerno gets the credential the way a client would, by calling your sign-up or login route. Where no such route exists, it seeds the credential directly in your datastore or identity provider, one of the cases direct access exists for.

Some entry points require Kerno to construct the credential itself. It reads the exact claims your verifier enforces, including the audience and issuer from your own configuration, and mints a properly signed JWT using the signing secret you supplied. This is real authentication. Your application verifies the token exactly as it verifies production traffic.

{% hint style="info" %}
**Kerno never keeps your secret values**. It records only the name of the environment variable each secret lives in, such as `DB_PASSWORD`. You put the value in `.kerno/.env`, which Kerno keeps out of git and reads when a test runs.
{% endhint %}


# Baseline Tests

Understand how Kerno generates test scenarios for your entry points, what a baseline is, and how both stay in sync as your code evolves.

A baseline test reflects how your app currently behaves. Kerno captures your entry point's current behavior as a test, then assesses every future code change against it. If a change alters how your entry point behaves, the baseline catches it.

### Test scenarios

Scenarios are the test cases Kerno generates for your entry points. Each one describes what state to set up, what request to send, what response to expect, and how to verify side effects. They live under `.kerno/scenarios/endpoints/` in your repository and are meant to be committed alongside your code.

Kerno also tests [user flows](/docs/core-concepts/user-flows), journeys that walk several entry points in order. Each flow gets one test, stored under `.kerno/scenarios/flows/`.

Every scenario is written twice, as a plain-English `.scenario.md` for you and an executable `.scenario.ts` for the sandbox. See [Anatomy of a Kerno test](/docs/references/anatomy-of-a-kerno-test) for how to read them.

{% hint style="info" %}
Kerno always writes scenarios in TypeScript, whatever language your application is written in. They reach your app over HTTP, so the language of your service does not matter.
{% endhint %}

A test scenario runs in four phases.

* **Arrange.** Set up any state the test needs, such as seeding a user, obtaining a token, or inserting a row.
* **Act.** Make the request to the entry point under test.
* **Assert.** Check the response and any observable side effects.
* **Clean up.** Tear down test data and close connections.

### What baseline test record

A baseline records what a correct response looks like. It captures the status code, every field in the response body, the response headers that describe behavior, and any state the entry point stored.

Kerno builds the baseline from your entry point's own data contract, the response types defined in your code. So it checks that every field holds the type and value your handler produces, deciding each expected value from its source:

* A value your handler always produces the same way, like a fixed status code, a constant string, or a field copied straight from the request, is recorded exactly. Any change to it shows up as a diff.
* A value generated fresh on every run, like an id, a token, or a timestamp, is recorded by its format. A new id on each run never counts as a diff.

The baseline also fixes the exact set of fields a response returns. A field the entry point did not return before shows up as a diff, which is how an accidentally exposed field, like an internal flag or a leaked password hash, gets caught.

For HTTP routes, the baseline also records headers such as `Content-Type`, `Cache-Control`, `Set-Cookie`, the CORS headers and the security headers, each with its current value or as absent. A header-only change, like a CORS policy switched to a wildcard, shows up as a diff. Headers that change on every run, like `Date` or a request id, are left out.

A baseline never records a leaked error detail, such as a stack trace, a file path, SQL text or an exception class name, as the expected error response. In security tests, each error response is also checked for the absence of those details.

### What Kerno tests

Kerno generates two kinds of tests. **Validation tests** check that your entry point behaves correctly, and **security tests** check that it resists the vulnerabilities that apply to it. Kerno generates validation tests, and adds security tests when you ask for them.

{% hint style="info" %}
Before a run, your agent may confirm with you which kinds of tests you want, the effort level, and whether to test white box or black box, unless you already said so. Kerno keeps no preferences of its own. Your agent can save your answer in its own instruction files, such as `CLAUDE.md` or `AGENTS.md`, so it stops asking.
{% endhint %}

#### **Validation tests**

For each entry point, Kerno generates validation scenarios across several categories of behavior, including:

<table><thead><tr><th width="197.994873046875">Category</th><th width="581.7186279296875">Description</th></tr></thead><tbody><tr><td><strong>Functional API Workflows</strong></td><td>Covers entry point behaviour, multi step request sequences, coordinated service interactions, and integration patterns across services.</td></tr><tr><td><strong>Contract &#x26; Schema Validation</strong></td><td>Checks request and response structures, data types, required fields, serialization rules, and version compatibility for the API.</td></tr><tr><td><strong>Error Handling &#x26; Resilience</strong></td><td>Validates status codes, error body formats, retry behaviour, backoff procedures, timeout handling, and controlled fallback behaviour.</td></tr><tr><td><strong>Authorization &#x26; Authentication</strong></td><td>Reviews token validation, role based access rules, permission scopes, session handling, and all credential related flows.</td></tr><tr><td><strong>Boundary &#x26; Edge Cases</strong></td><td>Examines payload size limits, pagination behaviour, null or empty values, malformed inputs, and constraint based validation.</td></tr><tr><td><strong>Data Integrity &#x26; Persistence</strong></td><td><p>Confirms data consistency, transaction behaviour, idempotent operations,</p><p>state handling, and enforcement of database rules.</p></td></tr></tbody></table>

### **Security tests**

Security tests baseline your entry points' security posture. Kerno captures how the entry point stands up to a specific class of attack today, so if a later change opens that vulnerability, the test starts failing and the regression is caught.

Kerno assesses which **OWASP API and Web Top 10 categories** plausibly apply to the entry point, given its inputs, its auth model, and what the surrounding code does, then writes tests only for the categories that fit. Common ones include **broken object-level and function-level authorization**, **injection**, **excessive data exposure**, **mass assignment**, and **server-side request forgery**.

### Testing MCP servers

Kerno tests each tool your MCP server exposes as its own entry point. It calls the tool with arguments and checks what comes back, the returned content, the structured output against its schema, and whether the call succeeded or returned an error.

For each tool, Kerno generates:

* **A happy-path call** with valid arguments, checked against the tool's output contract.
* **A case per required argument**, omitted or given the wrong type, expecting the tool to return an error.
* **Type and boundary cases** for constrained fields.
* **Side-effect checks** for tools that change state, which assert the change happened and then clean up.

The baseline records the result the tool returns and masks non-deterministic fields, so a fresh id or timestamp never counts as a diff. The `validation` and `security` tags apply as they do for HTTP, so a tool can carry both functional and OWASP security coverage.

To point Kerno at a tool, ask your agent to test it by name. `kerno_list_mcp_tools` lists what is testable, and tools also appear in your normal entry point list.

{% hint style="info" %}
Kerno connects to your MCP server over Streamable HTTP. Point your `sut_url` at the server's MCP path, for example `http://localhost:9300/mcp`. Discovery and planning work from your source with the server stopped. Implementing and running scenarios need it up, the same rule that applies to HTTP routes.
{% endhint %}

If your MCP server requires authentication, ask your agent to add the headers it expects:

```
My MCP server needs an Authorization header. Add it to Kerno's MCP_AUTH_HEADERS in .kerno/.env.
```

`MCP_AUTH_HEADERS` is a JSON map of header names to values, such as `{"Authorization": "Bearer <token>"}`. Kerno sends these headers only to the address configured for your app.

### Testing background consumers

Kerno also tests the background work your application runs on its own. A consumer is started by a queued job or a broker message instead of an HTTP request, so it has no caller identity and returns no response. Its observable outcome is a store write or a message it publishes in turn, and that outcome is what the baseline records.

Address a consumer by its trigger verb and its registered name.

| Trigger                                                          | `endpoint_method` | `endpoint_path`                                                                   |
| ---------------------------------------------------------------- | ----------------- | --------------------------------------------------------------------------------- |
| A Celery task, in a Python application                           | `TASK`            | the dotted name a worker registers it under, such as `orders.tasks.process_order` |
| A Graphile Worker job, in a TypeScript or JavaScript application | `TASK`            | the identifier the worker registers, such as `send_welcome_email`                 |
| A handler a broker message triggers                              | `MESSAGE`         | the handler's registered name, such as `payment-service.paymentRequestConsumer`   |

Celery tasks and Graphile jobs share the `TASK` verb. Kerno tells them apart by the language of the file it found the task in, and dispatches each the way its own worker expects, through the broker for Celery and through `add_job` for Graphile.

The path is the registered name, so it carries no leading slash and looks nothing like a URL.

Kerno discovers consumers from your source. For Celery it reads one entry point per decorated function, recording the queue each is dispatched on and the parameter list a dispatched message has to satisfy. For Graphile Worker it reads all three places a real application registers tasks, which are the crontab file, a task directory where the file name is the identifier, and a task-list object where the key is the identifier.

Consumers appear in your normal entry point list alongside HTTP routes and MCP tools.

A consumer scenario enqueues work that satisfies the handler's declared input, then asserts the side effect the handler produced. Because that outcome is a side effect rather than a response, verifying it usually needs the dependency access described in [Test environment](/docs/core-concepts/environment-setup).

### Effort levels

When Kerno creates baseline tests, the effort level sets how many tests it plans and how hard it tries to prove each one is solid, trading speed for rigor. It defaults to high, so you only lower it when you want a faster, lighter run. An entry point your team marked critical in the [Kerno portal](/docs/portal/overview) always runs at high.

* **Low.** Kerno plans a small representative set of at most 8 tests. Each test is written and run once. The fastest, cheapest signal.
* **Medium.** Kerno covers every parameter the handler reads, with one valid and one invalid test each. Each test runs twice in a row and both runs must pass. This proves the test is genuinely repeatable and exercises a second sample of randomized data.
* **High.** Everything medium does, plus a test for each value in the fixed sets your code accepts, such as an enum, up to 35 tests in all. A critique-and-repair pass reviews each finished test and sends it back until it holds up.

The second run is meaningful in itself. If a test passes once and fails the next, that is a real finding, either the test is not isolated or fresh random data hit an edge case worth knowing about.

### How Kerno ensures test quality

Tests run in Kerno's own sandbox against your live application, and every verdict comes from what the checks actually did there. Because the run is independent of your coding agent, the result reflects your application's real behavior.

Two mechanisms guard the tests themselves:

* **Tests are proven repeatable.** At `medium` and `high` effort each one runs twice in a row, so a test that only passed because of leftover state gets caught.
* **Tests are critiqued and repaired.** At `high` effort, each finished test goes through a review that checks it against what it was meant to test and sends it back for repair until it holds up. The review rejects checks weakened or rigged to pass no matter what, errors swallowed silently, tests that verify less than they were meant to, matching loose enough to pass on the wrong value, setup that claims state it never created, hardcoded identifiers where each run needs a fresh one, and teardown that claims to delete something without confirming it did.

### Reading test results

Each scenario reports one of four verdicts.

* **Passed.** The scenario ran and the entry point behaved as expected.
* **Failed.** The entry point did something the scenario did not expect.
* **Blocked.** The scenario could not run, usually because a dependency it needs is not configured. Blocked is neither a pass nor a fail; nothing was tested.
* **Not implemented.** Kerno could not produce a working scenario for this case. This is reported honestly rather than counted as a pass.

A passing scenario may also carry a **potential bug**, where Kerno found the entry point doing something unexpected, confirmed it against your source, and documented the real behaviour. See [Security testing](/docs/references/security-testing) for how to review and dismiss these.

#### Potential bugs

A baseline test can carry a **potential bug tag**. Kerno noticed the entry point doing something unexpected, confirmed it against your source code, and wrote the test to document that real behavior. The flag names what deviated and what a future change would signify. It names a suspected code location only when the run found evidence for one, and otherwise says the cause is unsettled.

A potential bug flags behavior worth your attention on a test that otherwise passes. If you review it and decide the behavior is known or intended, tell your agent to ignore it:

> "The `role` field Kerno flagged on `GET /users/:id` is intentional, we keep returning it for backwards compatibility. Have Kerno ignore that potential bug."

An ignore decision is recorded in Kerno's memory for that entry point, alongside the code revision you made it against. Later test results leave the flag out, though the test file keeps it until Kerno next updates that entry point's tests. There is currently no way to reverse it through Kerno, so read the flag carefully before you dismiss it. Clearing Kerno's cache for the whole workspace erases the decision, and the flag comes back. See [Memory and learning](/docs/core-concepts/memory-and-learning).


# Code Change Review

Understand how Kerno reviews your code changes and catches breaking changes and unintended side effects before they ship.

When you make a code change, Kerno re-runs your existing baseline tests against your running application and reports whether behaviour changed. If nothing changed, you get a clean run. If something did, Kerno shows you exactly what changed.

{% code expandable="true" %}

```mermaid
flowchart LR
    A[Code change] --> B[Sync workspace] --> C[Find impacted entry points] --> D[Run baseline tests]
    D --> E[No Diff Detected]
    D --> F[Diff Reported]
```

{% endcode %}

The review is targeted. Kerno traces your change to the entry points it affects and re-runs only their baseline tests, so each run stays focused on what you actually touched. It also reports the [user flows](/docs/core-concepts/user-flows) that cross those entry points, and your agent re-runs them too, so a change that breaks a journey is caught.

Each entry point and user flow Kerno reviews gets its own run report in the [Kerno portal](/docs/portal/overview). Your agent shares the link as soon as the review starts.

### Detecting diffs

A review starts from the blast radius of your change, every entry point the change can reach. Kerno runs those entry points' baseline tests and reports every diff in one run, so you see the full impact of your change.

Kerno finds the blast radius by comparing your working tree against HEAD, covering both staged and unstaged changes, and mapping the changed files to the entry points they reach. It follows the call graph across your whole codebase, so a change to a shared helper surfaces every entry point downstream of it.

See [Baseline tests](/docs/core-concepts/scenarios-and-baselines) for what a baseline checks and what counts as a diff.

### Review scope

Use a scope when asking Kerno which entry points need attention:

| Scope                     | Meaning                                                |
| ------------------------- | ------------------------------------------------------ |
| `all`                     | Every entry point in the selected application          |
| `changed`                 | Only entry points affected by your current git changes |
| `file:path/to/handler.ts` | Entry points defined in that source file               |
| `endpoint:METHOD /path`   | A single route, identified by HTTP method and path     |

See [Scopes](/docs/references/scopes) for the full reference.

### Responding to a diff

When Kerno reports a diff, it's up to you to decide whether it's a bug or an intended change.

* **It's a bug.** Your change broke something. Fix the code and validate again.
* **It's intentional.** The entry point is meant to behave this way now. Accept the diff, and Kerno updates the affected tests to match the new behaviour.

Sometimes a run ends without a verdict on your change. Kerno then says where the fault lies:

* **Your environment.** The app was unreachable, a configuration value was missing, or a dependency such as the database was down. Fix the setup and review again.
* **Kerno itself.** Kerno's sandbox failed to start or answer. Review again.

{% hint style="info" %}
When you accept an intended change, Kerno also finds any new logic your change introduced and adds baseline tests to cover it, so your coverage grows with your code.
{% endhint %}

{% code expandable="true" %}

```mermaid
flowchart LR
    A[Diff reported] --> B{Bug or intended?}
    B -->|Bug| C[Fix the code]
    C --> D[Review again]
    D --> A
    B -->|Intended| E[Update tests]
```

{% endcode %}


# User Flows

Understand user flows, the journeys through your app that Kerno tests end to end across several entry points.

A user flow is one thing a person does with your product, from start to finish. Signing up and creating a first project is a flow. So is adding items to a cart and checking out. Kerno tests a flow end to end, walking the entry points it crosses in order and checking that the outcome the user wanted actually happened.

Entry point tests check each way into your app on its own. A flow test checks that a whole journey across several of them works together, which is where integration bugs tend to hide.

```mermaid
flowchart LR
    A[Sign up] --> B[Create a project] --> C[Invite a teammate]
    C --> D[The teammate sees the project]
```

A flow can include any of your entry points, whether HTTP routes, MCP tools or background consumers.

### Mapping your flows

Your agent maps your flows with you, typically right after it [calibrates Kerno](/docs/core-concepts/environment-setup#calibration) to your repository. It asks which journeys matter most, keeps your own description of each one, and records the ones you confirm. Recording a flow runs nothing yet.

You can map more flows at any time:

```
Map the key user flows in this app with Kerno.
```

### Testing a flow

Ask your agent to test a journey:

```
Use Kerno to test the checkout flow.
```

Kerno plans one test that walks the steps in order and checks the outcome at the end. It always shows you the plan and waits for your approval before it writes the test. Once you approve, it writes the test and runs it against your application.

Flow tests always run at high effort, so a test must pass twice in a row before it is saved. They cover functional behaviour. [Security tests](/docs/references/security-testing) run on individual entry points.

The test is stored in your repository under `.kerno/scenarios/flows/`, in your application's folder, and you commit it alongside your code.

### After a code change

When you ask your agent to review a change, Kerno reports which flows the change touches, meaning every flow that crosses an entry point the change reaches. Your agent re-runs those flows along with the affected entry points' tests, so a change that breaks a journey is caught even when each entry point still passes on its own. See [Code change review](/docs/core-concepts/change-validation).

### Changing a flow

When a journey changes, tell your agent. It updates the flow, and Kerno plans its test again, showing you the new plan before anything runs.

### Where flows show up

Flows appear next to your entry points on the Coverage, Behaviour Drift and Runs screens of the [Kerno portal](/docs/portal/overview).

To test your first flow step by step, see [How to Test a User Flow](/docs/guides/test-a-user-flow).


# Custom Rules

Understand how to customize Kerno with custom rules, how to create them, and how they apply to the tests it generates.

Every codebase has rules that are obvious to your team and invisible in the source. Custom rules put those in your repository, so Kerno applies them on every run without you restating them.

Rules are guidance to the planner, not enforced constraints. Kerno weighs them against your source code, so a rule that contradicts verified behaviour never overrides it.

{% hint style="warning" %}
**Editing your config does not rewrite existing tests.** New rules shape the next generate for entry points that have no tests yet. For an entry point that already has tests, the existing tests are kept.

To apply new rules to an entry point that already has tests, ask your agent to generate it again with guidance to apply the new rules.
{% endhint %}

### Where rules live

Rules live in `.kerno/config.yaml`, under `test-generation`.

| File                 | Committed | Purpose                                                                     |
| -------------------- | --------- | --------------------------------------------------------------------------- |
| `.kerno/config.yaml` | yes       | Your rules and configuration, reviewed in pull requests like any other code |

`config.yaml` is committed, so a teammate who clones the repository inherits your rules, and changing one is a reviewable diff. Secret values stay out of it. Those live in `.kerno/.env`, which Kerno gitignores for you. See [Ports and file paths](/docs/references/ports-and-file-paths).

Calibration writes rules here too. When your agent [calibrates Kerno](/docs/core-concepts/environment-setup#calibration), it adds the confirmed facts, such as how to sign in and how to seed test data, to the workspace rules. Expect rules under `test-generation` that nobody on your team wrote by hand. They work like any other rule, and you can edit them the same way.

{% hint style="info" %}
Earlier versions of Kerno split rules across a committed `default.config.yaml` and a gitignored `config.yaml`. That split is gone. A workspace still holding a `default.config.yaml` has it folded into `config.yaml` automatically the first time the agent reads it, keeping the `config.yaml` value where both set the same field, and the old file is then removed.
{% endhint %}

### How rules are applied

Rules apply from broadest to most specific, and everything that matches is included, so a specific rule adds to the general ones.

```mermaid
flowchart LR
    A[Workspace context] --> B[Workspace entry point rules]
    B --> C[Application context]
    C --> D[Application entry point rules]
    D --> E[Merged guidance]
    F["Per-call guidance<br/>(wins on conflict)"] --> E
    E --> G[Planner]
```

For `POST /api/v1/payments` on the `payments-api` application, all four apply, in that order. For `GET /api/v1/invoices`, only the error-envelope and tenant rules apply; the payments-api rules do not.

The planner also reads two sources that rank below your rules:

* Context Kerno distills from your repository's documentation, such as `AGENTS.md`, `CLAUDE.md` and `README.md`. Text between `<!-- kerno instructions start -->` and `<!-- kerno instructions end -->` in those files is left out.
* Lessons Kerno learned from earlier runs. See [Memory and learning](/docs/core-concepts/memory-and-learning).

```yaml
# .kerno/config.yaml

test-generation:
  context: |
    Error responses use our standard envelope. On every failure, assert the
    body matches { error: { code, message } }, not only the status code.
  endpoints:
    "/api/v1/**": |
      Every v1 route is tenant-scoped. Always assert that a request
      authenticated as tenant A cannot read or modify tenant B's records.

applications:
  payments-api:
    test-generation:
      context: |
        Money amounts are integer minor units, never floats. Idempotency keys
        are required on every mutating request.
      endpoints:
        "POST /api/v1/payments": |
          Cover the duplicate-idempotency-key case explicitly: the second
          request must return the original payment, not create a second one.
```

#### Matching entry points

| Pattern                         | Matches                                  |
| ------------------------------- | ---------------------------------------- |
| `*`                             | Every entry point                        |
| `/api/users`                    | That exact path, any method              |
| `GET /api/users`                | That path and method only                |
| `/api/v1/*`                     | One path segment, e.g. `/api/v1/users`   |
| `/api/v1/**`                    | Any depth, e.g. `/api/v1/users/42/roles` |
| `regex:^/api/v[0-9]+/users/.+$` | Anything the expression matches          |


# Memory and Learning

Understand how Kerno remembers what it learns about your codebase and improves over time.

Memory and learning is how Kerno gets better as it works with your codebase and learns from your feedback.

```mermaid
flowchart LR
    A[Kerno needs an analysis] --> C{In memory and<br/>still current?}
    C -->|Yes| F[Reuse it]
    C -->|No| E[Work it out]
    E --> G[Store it, stamped with<br/>the current revision]
    U[A finding you review] --> G
    G --> F
```

Kerno learns in three ways:

* **From your code** – Kerno works out how your app is structured, which files serve which routes, and how each entry point authenticates. This analysis is slow, so Kerno does it once and reuses it, re-deriving only the parts a code change affects.
* **From running your tests** – While setting up and running tests, Kerno remembers what actually worked, like the steps needed to seed a user before a test can run. The next entry point in the same service reuses that instead of working it out again.
* **From your feedback** – When you dismiss a flagged issue, Kerno saves your decision and applies it on every later run. When you answer a question Kerno asks, like where to find a schema, Kerno reuses your answer for that entry point.

### Reviewing what Kerno learned

As Kerno tests your entry points, it writes down what it works out along the way. How a particular entry point authenticates, say, or what has to exist before a request will succeed. Kerno calls these lessons, and it reuses them when planning your next tests so it never works the same thing out twice.

Sometimes a lesson is wrong. Kerno shows you what it has learned so you can say so.

Ask your agent once a batch of tests has finished:

```
Show me what Kerno learned from that batch.
```

You get each lesson with what it says, which entry points it covers, and how many runs back it up. For each one you can:

* **Agree.** It matches what you saw. Kerno keeps using it and marks it as checked by you.
* **Amend.** The idea is right and the wording is off. Give Kerno better wording and it uses yours from then on.
* **Disagree.** It should not be there at all. Kerno stops using it and will not write it again.

The review also lists what Kerno's analyses read from your code, covering your repository's documentation, your build, your security setup, and the external services your app calls. Each analysis arrives as several records, one per part, and you judge each one the same way. These records come from reading your code, so they carry no run counts.

{% hint style="info" %}
Reviewing is optional and never holds up a run. Kerno carries on using a lesson you have not looked at, and labels it as unchecked while it does, so you can always tell what Kerno worked out on its own apart from what you confirmed.
{% endhint %}

### How Memory is Stored

Memory is kept as a git history alongside Kerno's working copy of your repository, on your machine. Each entry is a commit whose trailer records the code revision it was learned at:

```
kerno-scenarios | scenario memory: GET /users/me
---
refs/HEAD/d3e136b8ec06451739fe5428789268317b976472
```

Kerno also writes a copy of its memory records, such as lessons, into your repository at `.kerno/memory/`, one Markdown file per record with a header that shows whether it was reviewed, plus a generated `README.md`. The folder is meant to be committed, so your team can review what Kerno learned in pull requests, and Kerno removes any line in `.kerno/.gitignore` that would hide it. The copy leaves your machine only when you commit and push it. Kerno never reads it back.

Two consequences worth knowing:

* Memory persists across agent restarts. Clearing Kerno's cache for the whole workspace erases it, including the potential bugs you told Kerno to ignore, which then get flagged again. The copy in `.kerno/memory/` restores nothing.
* Because entries are commits, older versions remain, so a memory that was recomputed still has its previous value in history.

Every memory is stamped with the exact code revision it came from, so Kerno knows when a memory has gone stale and re-derives it.


# How to Configure the Kerno Test Environment

Learn how to connect Kerno to your running application, give it access to your database, and confirm the environment is ready to test.

### Introduction

Kerno tests your application by calling it over HTTP against your real stack, which means it needs your app running at an address it can reach. You connect the two through your coding agent, which saves the configuration, points Kerno at your app, and confirms Kerno can reach it.

In this guide, you will connect Kerno to your application, give it access to your database, and confirm the environment is ready to test.

### Prerequisites

Before you begin, you will need:

* Kerno installed and the agent running, with your coding agent connected over MCP. See the [Quickstart](/docs/getting-started/quickstart).
* Docker running on your machine.

### Step 1. Index your app

Tell your agent to set up Kerno for your project:

```
Set up Kerno for this project.
```

Kerno analyses your repository, indexes your codebase, and analyses its entry points. See [Codebase indexing](/docs/core-concepts/codebase-indexing).

Setup also adds Kerno's skills for your coding agent under `.agents/skills/`, linked from `.claude/skills/`, and lists them in your local git exclude file so they never show up as changes. Your agent may also add a short Kerno note to its own instructions file, such as `CLAUDE.md` or `AGENTS.md`.

If your repository contains more than one app, Kerno indexes all of them and your agent shows you the ones it found. Tell it which backend(s) you want to test.

Each one is configured separately, with its own URL and its own dependency access, so a monorepo's services can run on different ports and against different databases. They all live in the one `.kerno/config.yaml` at your repository root, keyed by application. See [Ports and file paths](/docs/references/ports-and-file-paths) for the shape.

### Step 2. Connecting your app

Kerno needs your app running before it can test it. Start it yourself, or ask your coding agent to start it for you, then point Kerno at it.

To have your agent start the app and connect Kerno in one go:

```
Start my app locally and set up the Kerno environment for it.
```

Or start the app yourself and tell your agent to connect Kerno to it:

```
Set up the Kerno environment for this app.
```

If your app runs somewhere else you can reach, such as a shared dev environment, give that URL instead.

Your agent saves the URL, and Kerno probes it from inside its sandbox to confirm it is reachable, so a wrong address is caught now instead of failing mid-test.

Kerno then reports whether the environment is ready. Your agent waits for `ready_for_endpoint_test` to report `true` before running any tests, and you can check at any time:

```
Check the Kerno environment status.
```

{% hint style="info" %}
Kerno reaches your app from inside its own sandbox and routes `localhost` to your machine, so an app bound to `127.0.0.1` works as is in most setups. Running containers do not mean Kerno is ready, so wait for the readiness signal before generating tests.
{% endhint %}

### Step 3. Giving Kerno database access (optional)

Kerno tests in one of two modes, depending on how much access you give it.

* **White box (default).** Kerno also connects to your dependencies, such as your database, cache, or queue. This lets it set up state and verify results your API does not expose, like seeding a record your API has no way to create, or checking that a stored password was hashed.
* **Black box.** Kerno talks to your app only over HTTP, like any external client, and works entirely through your own entry points.

Kerno defaults to white box, and uses direct dependency access once you have configured it. Any individual run can opt out by asking for black box. See [Testing modes](/docs/references/testing-modes).

To turn on white box, tell your agent to connect your dependencies. It works out the connection details from your project, so you do not need to give ports or URLs yourself:

```
Connect Postgres and Redis to Kerno.
```

Kerno also needs your database schema, which it derives from your source automatically for most projects. If it cannot, it pauses and asks you for a path rather than guessing.

Your credentials stay on your machine. `config.yaml` records only the name of the environment variable each secret lives in, such as `DB_PASSWORD`, and you put the real value in `.kerno/.env`, which Kerno gitignores for you. Kerno reads it when a test runs.

To keep Kerno away from a component your code talks to, such as a shared queue or a third-party service, tell your agent. Kerno records it as off limits and leaves it alone:

```
Kerno must never touch our Kafka cluster.
```

If your app calls other HTTP services, point Kerno at their OpenAPI specs, so tests match the real request and response shapes of those calls:

```
Our app calls the billing service. Its OpenAPI spec is at specs/billing.yaml. Add it to Kerno.
```

### Step 4. Calibrating Kerno

Once the environment is ready, your agent calibrates Kerno to your repository. This happens once per repository, and your agent starts it on its own. It works out how to sign in, how to seed test data, what makes writes fail and which infrastructure Kerno may touch. It checks each against your running app, then asks you to confirm them one at a time. Calibration records every piece of infrastructure your code talks to, so if you skipped Step 3, this is where you can give Kerno access.

If your agent doesn't start it, or you want to run it again, ask:

```
Calibrate Kerno for this repo.
```

Calibration also generates an environment healthcheck from your infrastructure. From then on, readiness depends on it too, so `ready_for_endpoint_test` stays `false` while your app or a dependency Kerno may access is failing its check. To re-run the check after a fix:

```
Check the Kerno environment status and refresh the healthcheck.
```

See [Calibration](/docs/core-concepts/environment-setup#calibration) for what it records and how to decline it.

### Step 5. Updating your configuration

Kerno stores your settings under `.kerno/` at your repository root:

```
.kerno/
├── config.yaml          # how Kerno reaches your app and its dependencies, committed
├── .env                 # the real secret values, gitignored
├── .gitignore           # keeps .env out of git
├── healthcheck.local.ts # the environment healthcheck from calibration, committed
├── memory/              # what Kerno has learned about your repository, committed
├── criticality.json     # entry points your team marked critical in the Portal, committed
└── scenarios/           # generated tests, committed
```

In a monorepo, each app keeps its own `scenarios/` and `criticality.json` in a `.kerno/` folder inside the app's directory.

`config.yaml` is committed, so your whole team shares one configuration and every change to it is reviewable in a pull request. It names the environment variable each secret uses. The values themselves live in `.kerno/.env`, which Kerno adds to `.kerno/.gitignore` for you. Your generated tests under `scenarios/` are committed as well.

You rarely edit these by hand. Your agent writes them and keeps them current as things change. If your app moves to a new port or a dependency changes, Kerno reports that it can no longer reach your app, and your agent works out the new details and saves them. You can also tell it directly:

```
My app moved to port 4000. Update the Kerno config.
```

### Next Steps

Your app is connected to Kerno, calibrated, and the environment reports ready. Kerno can now reach your app over HTTP, and if you gave it dependency access, it can set up and verify state directly as well.

Next, [capture a baseline](/docs/guides/capture-a-baseline). Kerno generates tests for your entry points and runs them against your app to record how it behaves today, which becomes the reference every future change is checked against.


# How to Create Baseline Tests for your Entry Points

Learn how to generate baseline tests for an entry point, review the plan, and run them to capture how your app behaves.

### Introduction

A baseline is a test that records how an entry point behaves against your running stack, and the reference point Kerno uses to detect changes later. Kerno captures one by analysing the entry point, writing a set of tests, and running them against your app to capture its actual responses.

In this guide, you will list your entry points, generate baseline tests for one and review the plan, then run them to capture the baseline.

### Prerequisites

Before you begin, you will need:

* Your application running, with Kerno pointed at it. See [How to Configure the Kerno Test Environment](/docs/guides/start-the-environment).
* The environment reporting `ready_for_endpoint_test`.

### Step 1. Listing your entry points

Ask your agent to list the entry points Kerno found:

```
Use Kerno to list the entry points in this app.
```

You get the entry points Kerno discovered, such as HTTP routes, background consumers and MCP tools, grouped by the file that defines them, each one showing whether it already has baseline tests.

{% hint style="warning" %}
In a large app, listing everything can return thousands of entry points. Narrow it with a scope to the part you care about, such as a single file or one entry point.
{% endhint %}

Narrow it to a single file:

```
Use Kerno to list the entry points in src/routes/users.ts.
```

Or a single route:

```
Use Kerno to list POST /users.
```

### Step 2. Generating baseline tests for an entry point

Pick an entry point and ask your agent to generate baseline tests for it:

```
Use Kerno to generate baseline tests for POST /users.
```

Before it starts, your agent confirms three choices with you, unless you have already stated them in the chat or in your agent's rules file:

* **Effort.** How hard Kerno works to prove each test is solid. `low` writes and runs each test once. `medium` runs it twice to prove it is repeatable. `high`, the default, adds an iterative pass that critiques and repairs each test until it holds up. Higher effort takes longer and produces sturdier tests. An entry point your team marked as critical in the Kerno Portal always runs at `high`, whatever you choose. See [Testing modes](/docs/references/testing-modes).
* **Testing mode.** White box, the default, lets Kerno use the dependency access you configured to set up and check state directly. Black box limits it to calling your app like any external client. See [Testing modes](/docs/references/testing-modes).
* **Test types.** Validation tests, security tests, or both. Validation, the default, checks that the entry point behaves correctly. Security assesses the entry point against the OWASP Top 10 and writes tests that probe for vulnerabilities. Asking for security on its own narrows the run to the happy path plus those vulnerability tests, so ask for both when you want your functional coverage as well. See [Security testing](/docs/references/security-testing).

With those set, Kerno analyses the entry point, works out how it authenticates and what it depends on, and plans a set of scenarios covering the happy path, error handling, edge cases, and authorization. It then stops to show you the plan.

#### 2.1 Covering multiple entry points

You can cover several entry points in one go. Ask your agent to work through a group of them, or a whole file:

```
Use Kerno to create baseline tests for all the entry points in my users API.
```

Your agent can cover them one at a time, approving each plan before moving to the next, or dispatch the whole group in a single batch. See [Testing many entry points in one call](/docs/references/kerno-mcp#testing-many-entry-points-in-one-call).

### Step 3. Reviewing the plan

The plan is Kerno's read of your entry point. It lists the test scenarios Kerno intends to write, each with the situation to set up, the request to send, and the response it expects. Reading it is the point of this step, because it shows you what Kerno believes your entry point does before it writes any code.

Approve the plan and Kerno starts implementing:

```
Looks good, go ahead.
```

Or send it back with changes. You can drop scenarios, add new ones Kerno missed, or ask for something specific:

```
Drop the pagination test and add one for a duplicate email.
```

```
Add a scenario that checks a non-admin cannot create a user.
```

Sending it back re-plans with your feedback and shows you a new plan. You can iterate as many times as you need.

{% hint style="info" %}
Changes you make to a plan apply to the current run. The feedback Kerno carries between runs is the decisions you make on results, like telling it a finding is intended so it stops flagging that entry point, along with what it learns about your entry point such as how to authenticate. See [Memory and learning](/docs/core-concepts/memory-and-learning).
{% endhint %}

Approving is what tells Kerno to write and implement the test scenarios.

### Step 4. Letting Kerno implement and run

Once you approve, Kerno writes each scenario and runs it against your application, repeating until it passes. Tes scenarios land under `.kerno/scenarios/endpoints/` as TypeScript files.

Before it runs anything, Kerno checks its **preconditions**, a set of readiness checks that confirm it can actually test the entry point. They verify the entry point is reachable, the intended user can authenticate, any required environment variables are set, and expected seed data exists. Whether your dependencies are up is the job of the [environment healthcheck](/docs/core-concepts/environment-setup#calibration). The checks read your data and change nothing, apart from one case. If you have not given Kerno a working test login, they may create a throwaway test user, through your app's signup route when there is one. What the entry point should do is left for the tests to decide.

If Kerno cannot satisfy the preconditions, it stops and tells you what it needs, such as a schema it could not derive or a credential, so you can provide it and run again.

{% hint style="info" %}
Kerno always writes test scenarios in TypeScript, whatever language your application is written in. They reach your app over HTTP, so the language of your service does not matter.
{% endhint %}

{% hint style="info" %}
Kerno writes two files per scenario. The `.scenario.md` is the one to read, since it describes the scenario in plain English. The `.scenario.ts` beside it is the executable version the sandbox runs. See [Anatomy of a Kerno test](/docs/references/anatomy-of-a-kerno-test).
{% endhint %}

### Step 5. Reading the result

When you are signed in to Kerno, your agent also gives you a link to the run's page in the [Kerno Portal](/docs/portal/overview), where you can follow it and read the results.

Kerno reports each scenario as one of:

* **passed** means the scenario ran and recorded the entry point's behaviour as your baseline.
* **failed** means Kerno wrote the scenario and ran it, but it did not pass. Kerno reports what went wrong so you can fix the cause and run it again.
* **blocked** means the scenario could not run because a precondition or capability it needs is missing, so nothing was tested. Often it needs direct database access that is not configured, and Kerno tells you which dependency to add to unblock it.

{% hint style="info" %}
You may occasionally see **not implemented**. This means Kerno could not produce a working scenario for a case and left the placeholder from the plan in place, rather than counting it as a pass. Nothing was tested.
{% endhint %}

A baseline captures how your entry point behaves today, including any bugs it currently has. Once your tests pass, read through the recorded responses and confirm they are what the entry point should return. If they look right, your baseline is ready. If something is wrong, fix your code and run the baseline again so it records the corrected behaviour.

#### 5.1 Potential bugs

A scenario that passed can also carry a **potential bug**. Kerno found the entry point doing something the plan did not expect, confirmed it against your source, and wrote the test to document the real behaviour. The note explains the root cause and what a future change would mean.

If the behaviour is known or intended, tell your agent to ignore it, and Kerno stops flagging it on that entry point:

```
The role field on GET /users/:id is intentional, we return it for backwards compatibility. 
Have Kerno ignore that potential bug.
```

Later runs, including reviews of your code changes, no longer report it. The test file keeps its note about the bug until you ask Kerno to update that entry point's tests. Clearing Kerno's whole cache forgets the decision, and the bug is reported again.

### Editing baseline tests

To change what a baseline test does, ask your agent to update it rather than editing the file by hand. Kerno rewrites the test and re-runs it against your app, so the baseline still reflects real behaviour.

```
Use Kerno to update the baseline tests for POST /users. 
The response no longer includes the role field.
```

### Next Steps

You now have baseline tests capturing how your entry points behave today. From here, Kerno re-runs them each time you change code and tells you what moved, so a regression surfaces the moment it appears.


# How to Review your Code Changes

Learn how to review a code change against your baseline tests, read the result, and update them when the change was intentional.

### Introduction

Reviewing your changes is the day-to-day loop. You change some code, Kerno re-runs the baseline tests for the affected entry points, and tells you whether their behaviour moved.

In this guide, you will trigger a review, read the result, and bring the tests along when a change was intentional.

### Prerequisites

Before you begin, you will need:

* Baseline tests already generated for the entry points you want to review. See [How to Create Baseline Tests for your Entry Points](/docs/guides/capture-a-baseline).
* Your application running with your changes applied.

### Step 1. Triggering a review

After changing code, ask your agent to review the entry points your change affected:

```
Use Kerno to validate the entry points and flows my changes affected.
```

Kerno maps your change through the call graph to every entry point it reaches, so a change to a shared helper is reviewed alongside the obvious ones. For each affected entry point, it re-runs the baseline tests already on disk against your current code and reports where behaviour moved. It also re-runs the [user flows](/docs/core-concepts/user-flows) that cross those entry points.

Kerno reviews the changes you have not committed yet, staged or unstaged, against your last commit. Run the review before you commit, and `git add` any new files first, since Kerno only sees files git tracks. See [Code Change Review](/docs/core-concepts/change-validation#detecting-diffs).

You can also review a single entry point:

```
Use Kerno to validate POST /users.
```

A review only re-runs tests that already exist. If an entry point has no tests yet, Kerno reports that it has none, so ask your agent to [create a baseline](/docs/guides/capture-a-baseline) for it first.

### Step 2. Reading the result

When a review finishes, Kerno re-runs each baseline test against your current code and checks whether the response still matches the one the baseline recorded. When you are signed in to Kerno, your agent also gives you a link to each review's page in the [Kerno Portal](/docs/portal/overview).

A diff is a difference between the response the baseline recorded and what your entry point returns now. Kerno matches values that legitimately change between runs, such as generated ids, tokens, and timestamps, by their shape, so a fresh id or timestamp never counts as a diff. A field that appears where the baseline had none does count, even without an explicit assertion for it, which is how an accidentally leaked field, like an internal flag or a password hash, gets caught. The same goes for response headers that describe behaviour, such as content type, caching, cookies, CORS and security headers, so a header that changes, appears or disappears is a diff too.

Each test comes back as one of four outcomes:

* **Passed.** No diff was detected. The behaviour this test covers matches the baseline, so your change left it unchanged.
* **Blocked.** The scenario never ran, usually because a dependency it needs is not configured. Nothing was tested. Kerno tells you what is missing, so you can ask your agent to fix it and run the review again. See [How to Configure the Kerno Test Environment](/docs/guides/start-the-environment).
* **Failed.** A diff was detected. The behaviour moved from the baseline. Kerno reports every assertion that did not match, each with a clue explaining what it was checking and a diff of expected against actual, so you see every change from one run.
* **Not implemented.** The test still holds the placeholder from its plan, because Kerno never finished writing it. Nothing was tested.

A diff does not always mean something is broken. It only tells you the behaviour changed, and it is up to you to decide:

* **The change is a bug.** Fix your code and review again.
* **The change is intentional.** Your entry point genuinely changed and the tests need to follow, so update the baseline (Step 3).

### Step 3. Updating baseline tests

Ask your agent to update when a diff is intentional, or when you have added logic that no test covers yet:

```
Use Kerno to update the tests for POST /users. 
The response no longer includes the role field.
```

Kerno reconciles the suite with your change:

* It re-baselines the tests that diffed, so their expectations match the new behaviour.
* It adds tests for new logic your change introduced that nothing covers yet.
* It keeps every test that still applies, and removes one only when you tell it that test is obsolete.

{% hint style="info" %}
When you add logic that none of the existing tests exercise, the review shows no diff. Ask Kerno to update the tests for that entry point and it adds coverage for the new logic:

```
Use Kerno to update the tests for POST /users.
```

{% endhint %}

### Conclusion

That is the loop. Change code, review the affected entry points, and update the tests when a change is intentional, so your baseline always reflects how your app should behave. From here, see [Custom Rules](/docs/core-concepts/custom-rules) to shape how Kerno tests, and [Memory and learning](/docs/core-concepts/memory-and-learning) to see how it improves over time.


# How to Test a User Flow

Learn how to map the key journeys through your app, test one end to end, and re-check it after a code change.

### Introduction

A user flow is a journey a person takes through your product, such as signing up and creating a first project. Kerno tests it end to end, across every entry point it crosses, and checks that the outcome actually happened. See [User flows](/docs/core-concepts/user-flows) for the concept.

In this guide, you will map your key flows, test one, and re-check it after a code change.

### Prerequisites

Before you begin, you will need:

* Kerno set up, with your environment ready and calibrated. See [How to Configure the Kerno Test Environment](/docs/guides/start-the-environment).
* Your application running.

### Step 1. Mapping your flows

Your agent usually maps your flows with you right after calibration. To map them, or to add more later, ask:

```
Map the key user flows in this app with Kerno.
```

Your agent asks which journeys matter most, then records the flows you confirm in your own words. Recording a flow runs nothing yet.

### Step 2. Testing a flow

Pick a journey and ask your agent to test it:

```
Use Kerno to test the checkout flow.
```

Kerno plans one test that walks the steps in order and checks the outcome at the end. If the journey needs an account or credential Kerno doesn't have yet, Kerno may ask you for it through your agent.

### Step 3. Approving the plan

Kerno always shows you the plan before it writes the test. Read it through and check that the steps and the expected outcome match the journey you meant, then approve it, or tell your agent what to change.

Once you approve, Kerno writes the test, checks that it passes twice in a row, and runs it against your application. If the journey breaks, the result shows the step where it failed.

### Step 4. Re-checking after a change

After you change code, ask your agent to review it:

```
Use Kerno to validate the entry points and flows my changes affected.
```

Kerno reports which flows your change touches, and your agent re-runs them along with the affected entry points' tests. When a journey changed on purpose, tell your agent, and Kerno plans the flow's test again for your approval.

### Next Steps

Your first flow is tested, and it is re-checked whenever you review a change that reaches it. Map the rest of your key journeys the same way, and see [How to Review your Code Changes](/docs/guides/validate-code-changes) for reading the results of a review.


# How to Customize Kerno with Custom Rules

Learn how to add a custom rule, apply it to your entry points, and edit or remove   rules as your team's conventions change.

### Introduction

Custom rules are standing guidance you keep in your repository, so Kerno applies them on every run without you restating them. They shape how the planner writes tests, and they travel with your code.

In this guide, you will add a rule, apply it to an entry point, and then edit and remove rules. For how rules are weighed, merged, and matched to entry points, see [Custom Rules](/docs/core-concepts/custom-rules).

### Prerequisites

Before you begin, you will need:

* [Kerno installed](/docs/getting-started/quickstart) and connected to your coding agent over MCP.
* Your application running, with the [Kerno test environment configured and ready.](/docs/guides/start-the-environment)

### Step 1. Add a rule

The quickest way to add a rule is to ask your agent:

```
Add a Kerno rule for POST /api/users: a duplicate email should return 409.
```

Your agent writes the rule into `.kerno/config.yaml` under `test-generation`. That file is committed, so your team inherits your rules. A `context` block applies to every entry point, and an `endpoints` map scopes guidance to specific routes:

```
test-generation:
  context: |
    ...
  endpoints:
    "POST /api/users": |
      ...
```

You can also edit this file by hand. The entry point pattern grammar is covered in [Custom Rules.](/docs/core-concepts/custom-rules)

### Step 2. Apply and verify

A new rule shapes the next `generate` for an entry point that has no tests yet. Ask your agent to generate that entry point:

```
Use Kerno to generate tests for POST /api/users.
```

When the plan comes back for review, the guidance shows up in the scenarios it proposes, here a scenario covering the duplicate-email case. Approve the plan to implement them.

### Step 3. Apply a rule to an entry point that already has tests

To bring an entry point that already has tests in line with a new rule, ask your agent to update its tests:

```
Use Kerno to update the tests for POST /api/users to follow the new duplicate-email rule.
```

Kerno keeps the tests that still apply, revises the ones the rule changes, and adds the ones it is missing. Editing the config on its own leaves existing tests as they are, see [Custom Rules](/docs/core-concepts/custom-rules). Generating again keeps the tests you already have, and generating with new one-off guidance replaces the whole suite with a fresh plan once you approve it.

### Step 4. Edit or remove a rule

Ask your agent to change or remove a rule:

```
Update the Kerno rule for POST /api/users to also assert the response includes the created user id.

Remove the Kerno rule for POST /api/users.
```

You can also edit `.kerno/config.yaml` by hand, change a rule's text to edit it, or delete its `context` block or `endpoints` entry to remove it.

Edits and removals behave like additions. They change what Kerno plans on the next run, and they leave existing tests untouched. Update the affected entry points' tests when you want the change to apply.

Where several rules match the same entry point, Kerno applies all of them, and guidance you give for a single run takes precedence when they conflict. For the full merge and precedence order, see [Custom Rules](/docs/core-concepts/custom-rules).

### Conclusion

That is the rule lifecycle. Add a rule, generate to apply it, and edit or remove it as your conventions change, updating the affected entry points' tests each time. From here, see [Custom Rules](/docs/core-concepts/custom-rules) for how rules are weighed and matched, [Testing modes](/docs/references/testing-modes) for per-run guidance, and [How to Review your Code Changes](/docs/guides/validate-code-changes) for the everyday validation loop.


# How to Run your Baseline Tests in CI

Learn how to replay your committed baseline tests against your application on every pull request, and read the result as a check.

### Introduction

The baseline tests Kerno writes are files in your repository. Once they are committed, anything that can run them can run them, including your CI.

This guide wires up the Kerno GitHub Action. On every pull request it replays the baseline tests you have committed against your running application and reports the result as a check, so behaviour that moved is caught where your team already reviews changes rather than on one developer's machine.

It needs **no Kerno account, no API key and no agent**. The action pulls one public image and runs the tests already in your repository. Nothing calls a language model, which is also why it is fast and free to run. To also follow each run in the Kerno portal, add an API key ([Step 6](#step-6-seeing-runs-in-the-kerno-portal)).

### Prerequisites

Before you begin, you will need:

* Baseline tests committed under `<app>/.kerno/scenarios`. See [How to Create Baseline Tests for your Entry Points](/docs/guides/capture-a-baseline).
* A way to start your application in CI, and a URL it answers on.
* A Linux runner with Docker. `ubuntu-latest` works as-is.

{% hint style="info" %}
Tests that read a database, mint tokens from a shared secret, or call a downstream service need values you would never commit. [Step 4](#step-4-tests-that-need-configuration) covers passing those in from secrets.
{% endhint %}

### Step 1. Adding the workflow

Create `.github/workflows/kerno.yaml`:

```yaml
name: kerno

on: pull_request

permissions:
  contents: read
  checks: write          # the reporter step creates a check run
  pull-requests: write   # the action comments its results on the pull request

jobs:
  kerno:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5

      # Your application. Start it however you already do — Kerno only connects to it.
      - run: docker compose up -d --wait
      - run: |
          for _ in $(seq 1 30); do
            curl -fsS -o /dev/null http://localhost:8080/health && exit 0
            sleep 2
          done
          echo "app did not become ready" >&2; exit 1

      - uses: kernoio/kerno-check@v1
        with:
          sut-url: http://localhost:8080

      - uses: mikepenz/action-junit-report@v6
        if: always()
        with:
          report_paths: kerno-junit.xml
```

Three things about that file are worth knowing:

* **You start your application; Kerno connects to it.** The action never starts, builds or tears down your app. Whatever you already use, whether Compose, a service container, or `./gradlew bootRun &`, keep using it.
* **Poll for readiness, not just for health.** `compose up --wait` waits for container health, which is not the same as your application serving requests. The loop above is worth copying.
* **`localhost` is fine.** The tests execute inside a container, where `localhost` would mean the container itself, so the action rewrites a `localhost` or `127.0.0.1` URL to reach your runner.

`@v1` follows every 1.x release. If you need a reference that never moves, pin a full version such as `@v1.2.2`, or a commit SHA. Versions before `v1.1.0` lack the `forward-env` and `apps` inputs this guide uses.

### Step 2. Reading the check

The action writes a summary to the job page:

```
⚠️ Kerno scenarios — my-service
30/47 passed — 17 skipped (never executed) against http://host.docker.internal:8080
```

A run with nothing failed and nothing skipped is marked ✅; one with skips is marked ⚠️ so they are visible at a glance, and a failure is marked ❌.

A failure is listed with the assertion that did not match and a diff of expected against actual, so you can see what moved without opening the logs. The JUnit report lands at `kerno-junit.xml`, which the reporter step turns into a check run with each test's result.

On a pull request, the action also posts its totals as a comment, and updates the same comment on every run. The comment and the check run need the `pull-requests: write` and `checks: write` permissions in the workflow above. Without them, they don't appear, and the action only logs a warning.

Each test comes back as one of the four verdicts described in [Baseline tests](/docs/core-concepts/scenarios-and-baselines#reading-test-results). JUnit has three states, so they map like this:

| Kerno verdict   | In the report | Fails the check |
| --------------- | ------------- | --------------- |
| Passed          | passed        | no              |
| Failed          | failed        | **yes**         |
| Blocked         | skipped       | no              |
| Not implemented | skipped       | no              |

**Skipped tests never fail the check, and they are always counted separately.** A blocked test is a known state, meaning a dependency it needs is not configured, and a not-implemented one is reported honestly rather than counted as a pass. That is why the summary reads `30/47 passed — 17 skipped` rather than `47 passed`: a suite that asserts nothing should not look like coverage.

### Step 3. Choosing what runs

By default the action discovers every `<app>/.kerno/scenarios` tree in your repository and replays all of them, which is usually what a monorepo wants.

To replay one application:

```yaml
      - uses: kernoio/kerno-check@v1
        with:
          sut-url: http://localhost:8080
          app-dir: services/orders
```

To replay a subset, filter on the path relative to the scenarios directory:

```yaml
      - uses: kernoio/kerno-check@v1
        with:
          sut-url: http://localhost:8080
          scenarios: endpoints/GET/**
```

`*` stays within a path segment and `**` crosses them.

### Step 4. Tests that need configuration

Tests that read or seed a database, mint tokens from a shared secret, or call a downstream service need values that must not live in your repository. Name them in `forward-env` and supply them from secrets:

```yaml
      - uses: kernoio/kerno-check@v1
        with:
          sut-url: http://localhost:8080
          forward-env: |
            DATABASE_URL
            JWT_SECRET
        env:
          DATABASE_URL: ${{ secrets.DATABASE_URL }}
          JWT_SECRET: ${{ secrets.JWT_SECRET }}
```

Two properties worth relying on:

* **Only the names you list are forwarded.** The rest of the runner's environment does not reach the container.
* **A name with no value stops the run**, before the container starts, and says which variable is missing. So does a name that isn't a valid variable name, or one of the runner's own variables (`PATH`, `HOME`, `NODE_OPTIONS`, `LD_PRELOAD` and `LD_LIBRARY_PATH`). That matters more than it sounds: a test that needed `DATABASE_URL` and did not get it fails at its database step while every HTTP assertion around it still passes, which reads as a partial success rather than a configuration mistake.

### Step 5. A monorepo with services on different ports

`sut-url` gives every application the same address. When your services listen on different ports, map each one instead:

```yaml
      - uses: kernoio/kerno-check@v1
        with:
          apps: |
            services/orders=http://localhost:8080
            services/billing=http://localhost:8081
          forward-env: |
            DATABASE_URL
        env:
          DATABASE_URL: ${{ secrets.DATABASE_URL }}
```

Each directory is the one containing `.kerno`, relative to the repository root, and each application is replayed against its own URL. `apps` replaces both `sut-url` and `app-dir`, and setting `apps` with either one stops the run with a configuration error.

One report is written per application, so `report-path` is a **directory** in this mode, and the reporter step takes a glob:

```yaml
      - uses: kernoio/kerno-check@v1
        with:
          apps: |
            services/orders=http://localhost:8080
            services/billing=http://localhost:8081
          report-path: kerno-reports

      - uses: mikepenz/action-junit-report@v6
        if: always()
        with:
          report_paths: kerno-reports/*.xml
```

The counts are summed across every application, so one failing test in one service fails the check.

### Step 6. Seeing runs in the Kerno portal

Add your Kerno API key and organization, and the action opens a run in the [Kerno portal](/docs/portal/runs) for each entry point it replays. The pull request comment then lists each run with a link to its report:

```yaml
      - uses: kernoio/kerno-check@v1
        with:
          sut-url: http://localhost:8080
          api-key: ${{ secrets.KERNO_API_KEY }}
          organization-id: ${{ vars.KERNO_ORGANIZATION_ID }}
```

Create the key under **Settings → API Key** in the portal (see [Integrations and API keys](/docs/portal/integrations)), store it as a repository secret, and set both inputs together. Setting only one of them stops the run with a configuration error. A run link that can't be opened never fails the check.

The summary also says how many entry points your team marked critical in the portal, and where that count came from. With credentials, the action reads it from the portal. Without them, it reads the `criticality.json` committed under each application's `.kerno/` directory.

### Step 7. Tracking your default branch

To let the portal follow how your default branch's tests grow over time, add a second workflow that runs `mode: sync` after every merge:

```yaml
name: kerno-sync

on:
  push:
    branches: [main]
    paths: ['**/.kerno/**']
  workflow_dispatch:

permissions:
  contents: read

jobs:
  sync:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - uses: kernoio/kerno-check@v1
        with:
          mode: sync
          api-key: ${{ secrets.KERNO_API_KEY }}
          organization-id: ${{ vars.KERNO_ORGANIZATION_ID }}
```

`sync` runs no tests and needs no running application. It sends the tests committed on your default branch to Kerno as that commit's snapshot, and lists in the job summary which entry points and tests the merge added or removed. It records only your default branch. Run it once by hand after you add it, to send the first snapshot.

### What this does not do

* **It does not create or update baseline tests.** That happens on a developer's machine, where the proposed tests can be reviewed before they are committed. CI replays what is already in the repository and nothing else.
* **It does not start your application.** You start it and pass `sut-url`. Kerno connects to a system under test; it never manages one.
* **It only forwards the environment variables you name.** Nothing else from the runner reaches your tests, and a name with no value stops the run rather than being quietly dropped.

### Reference

**Inputs**

| Input             | Required      | Default             |                                                                                                     |
| ----------------- | ------------- | ------------------- | --------------------------------------------------------------------------------------------------- |
| `mode`            | no            | `replay`            | `replay` runs your tests. `sync` sends your default branch's tests to the portal instead (Step 7).  |
| `sut-url`         | unless `apps` |                     | Base URL of your running application. A `localhost` URL is rewritten so the container can reach it. |
| `apps`            | no            |                     | One `<dir>=<url>` per line, replaying each application against its own URL. Replaces `sut-url`.     |
| `forward-env`     | no            |                     | Environment variable names to pass to your tests, one per line, valued from the step's own `env:`.  |
| `app-dir`         | no            | *(repository root)* | Replay one application's tests. Unset discovers every `<app>/.kerno/scenarios` tree.                |
| `scenarios`       | no            | *(all)*             | Glob filter on the path relative to the scenarios directory.                                        |
| `image`           | no            | *(pinned digest)*   | The runner image. Pinned by digest so a given version of the action always runs the same code.      |
| `report-path`     | no            | `kerno-junit.xml`   | Where the JUnit report lands. A directory when `apps` is used.                                      |
| `fail-on-failure` | no            | `true`              | Set `false` to report without gating.                                                               |
| `api-key`         | no            |                     | Your Kerno API key. Set it with `organization-id` to open portal runs, and for `mode: sync`.        |
| `organization-id` | no            |                     | Your Kerno organization. Set it with `api-key`.                                                     |

**Outputs**: `junit-path`, `total`, `passed`, `failed`, `skipped`, and `portal-run-urls` (one portal run link per line when `api-key` is set).

**Exit codes**

| Code |                                                                                                                                                                                                                                                                                                 |
| ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `0`  | Nothing failed. Tests may have been skipped, so check the counts.                                                                                                                                                                                                                               |
| `1`  | At least one test failed.                                                                                                                                                                                                                                                                       |
| `2`  | Configuration error: no tests found, an unparseable or empty report, or an input problem such as a missing URL, a malformed `apps` line, a missing scenarios directory, conflicting inputs, a bad `forward-env` name, or `api-key` without `organization-id`. Never reported as a test failure. |
| `3`  | The runner could not start. A broken container must not look like a failing test.                                                                                                                                                                                                               |

{% hint style="info" %}
The action gives the runner the `NET_ADMIN` capability so Kerno can intercept outbound HTTPS from your application. That is how a test can exercise a path that calls a third-party API without that API being reachable from CI.
{% endhint %}

### Conclusion

Your committed baseline tests now run on every pull request, and their exit code is your gate. Tests are still authored and reviewed locally, covered in [How to Review your Code Changes](/docs/guides/validate-code-changes), and CI is where the whole team finds out when behaviour moves.


# Kerno CLI

The Kerno CLI manages the Kerno agent on your machine and sets up MCP for your coding agent.

### Install

The CLI is published as [`@kerno/cli`](https://www.npmjs.com/package/@kerno/cli) on npm.

```bash
npm install -g @kerno/cli
```

Or run without installing.

```bash
npx @kerno/cli <command>
```

Requires Node.js 18 or higher. For the full setup walkthrough, see the [Quickstart](/docs/getting-started/quickstart).

### Commands

| Command             | Description                                                                                                                                                                                                  |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `kerno init`        | Start the Kerno agent and print the MCP server configuration for your coding tool. The main entry point.                                                                                                     |
| `kerno stop`        | Stop the Kerno agent. Cancels in-flight tests and stops application services first.                                                                                                                          |
| `kerno reset`       | Stop applications, clear workspace caches, and restart the agent fresh. The escape hatch when state is stuck.                                                                                                |
| `kerno status`      | Show current status.                                                                                                                                                                                         |
| `kerno doctor`      | Diagnose orphan agent processes and inconsistent state files. Use `--clean` to apply fixes.                                                                                                                  |
| `kerno logs`        | Stream workspace logs from the running agent. Use `-n <count>` to set the number of log lines (default 40).                                                                                                  |
| `kerno export-logs` | Export Kerno agent logs (workspace + terminal output) to a zip file. Requires `-o <path>` for the destination.                                                                                               |
| `kerno login`       | Log in to Kerno via your browser. With `--api-key <key>`, log in without a browser and keep the key for later commands. Add `--org <id>` when the key belongs to several organizations.                      |
| `kerno logout`      | Log out from Kerno.                                                                                                                                                                                          |
| `kerno update`      | Update the Kerno CLI to the latest version from npm. Works for a global `npm install -g` install. Other installs print the command to update with. Every other command warns when a newer version is on npm. |
| `kerno uninstall`   | Stop the agent and remove installed binaries.                                                                                                                                                                |

Run `kerno` with no arguments to launch the interactive shell.

### Global options

| Option                   | Description                                                    |
| ------------------------ | -------------------------------------------------------------- |
| `-w, --workspace <path>` | Workspace path. Defaults to the current directory.             |
| `-v, --verbose`          | Show Kerno diagnostic logs.                                    |
| `--force-update`         | Force re-download of the agent. Applies to `init` and `reset`. |
| `-V, --version`          | Show CLI version.                                              |
| `-h, --help`             | Show help.                                                     |

### `kerno init` options

| Option           | Description                                                                                                                                                                                                        |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--headless`     | Skip the interactive UI. Enabled automatically when stdin is not a TTY, such as inside a coding agent.                                                                                                             |
| `--force-switch` | In headless mode, stop an agent bound to a different workspace and rebind to this one.                                                                                                                             |
| `--org <id>`     | Organization to use for this workspace, by name or id. Recorded, so you are not asked again. A name must match exactly one of your organizations. With an API key, runs are recorded under the key's organization. |
| `--off`          | Turn Kerno off for this workspace and exit.                                                                                                                                                                        |
| `--on`           | Clear a previous `--off` for this workspace, so Kerno asks again on the next start.                                                                                                                                |

One agent serves several workspaces at once. Every MCP tool takes a `workspace_path`, and that argument selects which workspace the call targets. The agent accepts the root `kerno init` started it for and any directory nested under that root, so the applications in a monorepo are separate workspaces the same agent holds at the same time. It keeps up to 5 resident, which you can raise with the `KERNO_WORKSPACE_REGISTRY_MAX_WORKSPACES` environment variable.

Running `kerno init` from any other directory, including one nested under that root, rebinds the agent there. In a terminal it asks first. In headless mode it fails unless you pass `--force-switch`.

{% hint style="info" %}
If rebinding or restarting the agent moves Kerno to a different MCP port, update your coding tool with the URL `kerno init` prints.
{% endhint %}

### Environment variables

| Variable        | What it does                                                                                                   |
| --------------- | -------------------------------------------------------------------------------------------------------------- |
| `KERNO_API_KEY` | Signs in without a browser, for CI or a remote shell. Issue the key from **Settings → API Key** in the portal. |
| `KERNO_ORG_ID`  | Picks the organization when your API key belongs to more than one.                                             |

```bash
export KERNO_API_KEY=<your key>
kerno init -w /absolute/path/to/your/repo
```

Browser login needs a person in front of it, so use a key wherever that is not possible. See [Integrations and API keys](/docs/portal/integrations). The agent's port variables are covered in [Ports and file paths](/docs/references/ports-and-file-paths).

### Running inside a coding agent

The bare `kerno` shell renders an interactive terminal UI, which does not work in a non-interactive agent shell. `kerno status` works there. It runs its checks, prints them, and exits non-zero when one fails. `kerno --version`, `docker info`, and the `kerno_healthcheck` MCP tool also work from inside a coding agent.

### Examples

Start the agent and print MCP config.

```bash
cd /path/to/your/project
kerno init
```

Point the agent at a repository from anywhere.

```bash
kerno init -w /absolute/path/to/your/repo
```

Move the agent to a different repository.

```bash
kerno init -w /absolute/path/to/other/repo --force-switch
```

Fix orphan agent processes.

```bash
kerno doctor --clean
```

Clear workspace caches and restart cleanly.

```bash
kerno reset
```

Stream the most recent 100 log lines.

```bash
kerno logs -n 100
```

Export logs for support.

```bash
kerno export-logs -o ./kerno-logs.zip
```


# Kerno MCP

Kerno's MCP tools are the API your coding agent uses to drive Kerno.

### Async job pattern

Entry point tests, flow tests, validation runs and batches are async, because the underlying work can take minutes. Async tools return a `job_id` and a `kind` identifying the operation. Call `kerno_job` to wait for the result. `kerno_environment_setup` is synchronous, and returns a `job_id` only when it starts a background `healthcheck_author` job.

MCP hosts typically enforce a \~60-second timeout on a single tool call. `kerno_job` with `wait=true` blocks for at most 15 seconds, then returns the current snapshot with `wait_timed_out`, so call it again. Use `wait=false` for one immediate snapshot, block on the resource with `kerno_get_state` or `kerno_await_state`, or poll `GET /mcp/jobs` over HTTP.

When a job finishes, `kerno_job` reports `healthy`, `failed` or `cancelled`, or `completed` for a batch, and the result is also available via `GET /mcp/jobs`. A job parked at a feedback gate reports `needs_user_feedback` and resumes once the gate is answered. The last 50 finished jobs stay readable.

To cancel in-flight work, call `kerno_cancel` with a `job_id`, or with a `resource_id` or `batch_id` to stop one batch member or a whole batch. Cancellation is fire-and-forget, and the next `kerno_job` call returns `status: cancelled`.

{% hint style="info" %}
**Feedback gates show on the job.** When an entry point test pauses for plan approval or an answer, `kerno_job` reports `needs_user_feedback` with the prompt in `summary`, and `GET /mcp/jobs` reports the same status. Fetch the plan or question with `kerno_feedback_pending`, then answer with `kerno_feedback_answer`.
{% endhint %}

### Tools

#### Workspace management

| Tool                    | Description                                                                                                                                                                                                                                                                                            |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `kerno_sync_workspace`  | Sync Kerno's view of the workspace with the current state of code on disk. Cancels running Kerno operations, takes a fresh snapshot, and re-runs workspace analysis. Call this after making code changes.                                                                                              |
| `kerno_list_workspaces` | List the workspaces Kerno is managing, including snapshot state (branch, commit, uncommitted changes) and active jobs.                                                                                                                                                                                 |
| `kerno_clear_cache`     | Clear Kerno's cached state. Without `application_id`, it resets the whole workspace, including running jobs and learned memory, then re-runs analysis. Calibration declines are kept unless you pass `drop_declines: true`. With `application_id`, it resets only that application's test environment. |
| `kerno_list_mcp_tools`  | Introspect the MCP tools, resources and prompts your MCP server exposes. Pass `mcp_url` to query the running server and record what it advertises. Omit it to list what Kerno has recorded, detected from source the first time.                                                                       |

#### Discovery

| Tool                     | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_healthcheck`      | Check agent health, Docker, git, authentication, and codegraph. `workspace_path` is optional; defaults to the agent workspace.                                                                                                                                                                                                                                                                                                                                                                                                     |
| `kerno_get_applications` | Analyze the workspace and list applications Kerno detected, split into supported and unsupported, with each app's current config summary and recommended next action.                                                                                                                                                                                                                                                                                                                                                              |
| `kerno_list_endpoints`   | Browse what analysis found, grouped by source file, with which entry points already have tests. Covers HTTP routes, background consumers and the MCP primitives found in source, plus routes Kerno located on demand for a test. Tools a running MCP server advertises are listed by `kerno_list_mcp_tools`. Requires a `scope`, and can be filtered by app, coverage, method, protocol or file. With scope `changed`, it also returns the `impactedFlows` to re-run. The first call can take a few minutes on a large repository. |
| `kerno_graph`            | Return the workspace as a graph of applications, their entry points, the tests and flows that cover them, and their downstream services. Read-only, and returns in seconds. Hosts that support MCP Apps show it as an interactive graph that highlights what a running test touches.                                                                                                                                                                                                                                               |

#### Environment

| Tool                       | Description                                                                                                                                                                                                                                                                                                                                                                                          |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_save_config`        | Save how to reach your applications, one `applications` entry per app with its `target_environment` (`local` or `remote`) and `sut_url`, the URL as reached from your machine. An entry can also declare datastore access, downstream HTTP services, and which infrastructure Kerno may touch. Kerno probes the URL before saving, and saving an app starts its security analysis in the background. |
| `kerno_environment_setup`  | Probe the application and prepare it for entry point testing. Synchronous. Accepts an optional `sut_url` to save first, and returns `ready_for_endpoint_test` with the next step.                                                                                                                                                                                                                    |
| `kerno_environment_status` | Report readiness. Check `ready_for_endpoint_test` before running a test. Also reports the latest per-component health check, and `refresh_health: true` runs it now.                                                                                                                                                                                                                                 |

{% hint style="warning" %}
**Keep secret values in `.kerno/.env`.** `kerno_save_config` writes its blocks to `.kerno/config.yaml`, which is meant to be committed. In the block, declare only the variable name, with a placeholder value. Put the real value in `.kerno/.env`, which Kerno gitignores and merges in at higher precedence.
{% endhint %}

Kerno does not build or run your application. You start it yourself, then tell Kerno where it is. See [Environment Setup](/docs/core-concepts/environment-setup).

#### Entry point testing

| Tool                         | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_endpoint_test`        | Async. Generate or update tests for a single entry point. To re-run existing tests, use `kerno_validate`. Returns `job_id`, kind `endpoint_test`.                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `kerno_endpoint_test_many`   | Dispatches `kerno_endpoint_test` with `type` `generate` or `update` across many entry points in one call. Returns at once with a `batch_id` and one row per member, and runs as a job you can follow or cancel.                                                                                                                                                                                                                                                                                                                                                                  |
| `kerno_flow_test`            | Async. Plans, implements and runs one user-flow test, an ordered walk across several entry points. Takes `app`, a kebab-case `flow_id`, the user's own `flow_description`, and the `endpoints` in the order the user hits them. The test is stored under `.kerno/scenarios/flows/<flow_id>/`, and running the same `flow_id` again revises it. Returns `job_id`, kind `endpoint_test`.                                                                                                                                                                                           |
| `kerno_save_flows`           | Records the user flows you confirmed, without running them. Takes `app` and a `flows` list in the same shape as `kerno_flow_test`. Each flow comes back `saved` or `rejected` with the reason. A flow over HTTP routes that analysis has not found is still saved.                                                                                                                                                                                                                                                                                                               |
| `kerno_validate`             | Async. Re-runs the tests already on disk against your live stack and compares them with their baselines. Pass `endpoints` and `flows` to choose what to run. With neither, it validates every known entry point, and any entry point without tests reports "No scenarios found". A flow runs only when named in `flows`. Returns `job_id`, kind `validate`, with a Portal link for each entry point or flow. Each result carries a `fault_class` naming which side is at fault, such as `bad_environment`, and a `fault_reason` when the run stopped before any test was judged. |
| `kerno_ignore_potential_bug` | Record that a reported potential bug is known or intended. Later run reports for that entry point, `kerno_validate` included, leave it out. The test file keeps its flag until a `type: update` run rewrites it, and a full `kerno_clear_cache` discards the decision.                                                                                                                                                                                                                                                                                                           |

`kerno_endpoint_test` arguments:

| Argument                  | Required | Values                                                                                                                                                                                                             |
| ------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `workspace_path`          | yes      | absolute path                                                                                                                                                                                                      |
| `app`                     | yes      | application id from `kerno_get_applications`                                                                                                                                                                       |
| `endpoint_method`         | yes      | e.g. `GET`. For an MCP primitive, pass `TOOL`, `RESOURCE` or `PROMPT`. For a background consumer, pass `TASK` or `MESSAGE`                                                                                         |
| `endpoint_path`           | yes      | e.g. `/api/users`. A route analysis didn't list is located from its handler before the run, and refused only if no handler is found. For an MCP primitive, pass its name. For a consumer, pass its registered name |
| `type`                    | yes      | `generate` \| `update`                                                                                                                                                                                             |
| `effort`                  | no       | `low` \| `medium` \| `high`, default `high`. An entry point marked critical in the Portal always runs `high`                                                                                                       |
| `box_testing_strategy`    | no       | `black_box` \| `white_box`, default `white_box`                                                                                                                                                                    |
| `tags`                    | no       | `validation` (default) and/or `security`                                                                                                                                                                           |
| `test_generation_context` | no       | free-text guidance for this entry point                                                                                                                                                                            |
| `scenario_ids`            | no       | target specific scenarios only                                                                                                                                                                                     |
| `interactive`             | no       | boolean, default `false`                                                                                                                                                                                           |

See [Testing modes](/docs/references/testing-modes) for what each of these changes.

{% hint style="info" %}
**Testing MCP tools.** `kerno_endpoint_test` also tests the tools your MCP server exposes. Address a tool with `endpoint_method: TOOL` and `endpoint_path` set to the tool's name. Discover what is testable with `kerno_list_mcp_tools`. Kerno reaches the server over Streamable HTTP at the MCP path in your `sut_url`. See [Baseline tests](/docs/core-concepts/scenarios-and-baselines#testing-mcp-servers).
{% endhint %}

The launch response carries the `job_id`, a `monitoring_command` to block on until the plan needs your approval or the run finishes, and a `portal_run_url` for the run's Portal page. Its `resolved_intent` shows which values Kerno used and which it defaulted, and `effort_forced_by` appears when a critical entry point raised the effort.

#### Testing many entry points in one call

`kerno_endpoint_test_many` applies the same dials to a list of entry points. It returns straight away with a `batch_id`, a `job_id` and one row per member, each with the `resource_id` to quote back verbatim. To stop the batch, pass either id to `kerno_cancel`.

| Argument         | Required | Values                                                                                         |
| ---------------- | -------- | ---------------------------------------------------------------------------------------------- |
| `workspace_path` | yes      | absolute path                                                                                  |
| `app`            | yes      | application id                                                                                 |
| `endpoints`      | yes      | the entry points to cover                                                                      |
| `type`           | yes      | `generate` \| `update`                                                                         |
| `plan_approval`  | no       | `auto` (default) \| `gate`                                                                     |
| `max_in_flight`  | no       | caps how many members run at once, default 16. A member parked at its plan gate keeps its slot |

`effort`, `box_testing_strategy`, `tags` and `test_generation_context` apply to every member.

Under `plan_approval: auto` the plan gate never fires, so there is nothing to poll for. Under `gate`, members pause at the real plan gate and you answer each one with `answer_feedback_request`. Find the waiting members by blocking on `kerno_await_state` with the `batch_id`.

Each entry point is validated as it is dispatched. One that cannot be tested comes back with `admission` set to `rejected` along with the reason, and the rest of the batch carries on. Members over the `max_in_flight` cap come back `queued` and are promoted automatically.

This tool takes an explicit entry point list rather than a [scope](/docs/references/scopes).

#### Feedback gate

When Kerno needs input, such as a plan approval, a planner question, or missing credentials, it opens a feedback request on the relevant resource.

| Tool                      | Description                                                                                                                                                                                         |
| ------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_feedback_pending`  | Read-only list of the open feedback requests for one app, such as a plan awaiting approval, each with its `request_id`, `prompt` and `resource_id`. Returns `status: none` when nothing is waiting. |
| `kerno_feedback_answer`   | Answer by `workspace_path` + `app` + `request_id`. Payload is free-text `{"answer":"..."}` or approval `{"approved":true}` / `{"approved":false,"reason":"..."}`. Returns `accepted` once queued.   |
| `answer_feedback_request` | Answer by `resource_id` (the `.../feedback` subresource) + `request_id`. Same payload shape. Use this when you already have the `resource_id` from `kerno_get_state`.                               |

`kerno_feedback_answer` and `answer_feedback_request` reach the same action, so use whichever is more convenient given the context you already have.

#### Lesson review

Kerno writes down what it learns from your test runs and from analysing your code, and reuses it when planning later runs. These two tools let your agent see those records and judge them.

| Tool                    | Description                                                                                                                                                                                                                                         |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_lessons_pending` | Lists the records nobody has reviewed yet, each lesson with the entry points it covers and the runs that back it, and each analysis finding with the document it came from. `pending_total` is the full backlog, which can be larger than one page. |
| `kerno_review_lessons`  | Records a verdict on each record. `agree` keeps it and marks it checked. `amend` replaces its wording with yours. `disagree` retires a lesson for good and sets an analysis finding aside. Send every verdict in one call.                          |

Call these once a batch of entry point tests has finished, before starting the next one.

Kerno carries on using a lesson nobody has reviewed, and labels it as unchecked while it does. Reviewing never holds up a run. See [Memory and learning](/docs/core-concepts/memory-and-learning).

#### State plane

These tools let your agent read and watch the state of Kerno resources without polling a job.

| Tool                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_get_state`   | Read a resource's current state by `resource_id`. Returns `{resource_id, state}` where `state.status` is the variant tag, or `state=null` if the resource has never started. Pass `until_status` to long-poll until the status enters that set. `wait_timeout_ms` then defaults to 15000 ms, which is also the maximum, and a longer value is clamped. Without `until_status` the call returns immediately. `timed_out=true` is a normal outcome, so re-issue with the same `until_status`. Waiting states surface `open_feedback {resource_id, request_id, prompt}`.                                                                                         |
| `kerno_list_state`  | List every started resource under an optional `prefix`. Returns `{resources: [{resource_id, state}], timed_out}`. Pass `until_status` to wait until any of them enters that set, up to 15000 ms.                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| `kerno_await_state` | Wait for a batch of resources to reach a terminal status in one call. Name the set either by the `resource_id`s you dispatched or by the `batch_id` `kerno_endpoint_test_many` returned. `satisfy="all"` (default) returns once every id's status is in `until_status`. `satisfy="any_new"` returns as soon as an id that was already outside `until_status` enters it, so you can consume completions one at a time. Returns `{resources: [{resource_id, state}], timed_out}`. Use this in place of a polling loop. `any_new` is edge-triggered against the board as the call begins, so use `kerno_poll_events` when you need a cursor you can resume from. |
| `kerno_poll_events` | Replay domain events since a cursor. Returns `{stream_epoch, next_cursor, gap, events: [{seq, resource_id, event}]}` oldest-first. Start at `cursor=0`, then pass `next_cursor` back verbatim. If `stream_epoch` changes, the agent restarted, so reset to 0. If `gap=true`, events fell off the buffer, so re-read state. Accepts `prefix`, `max_events` (default 256) and `wait_timeout_ms` (maximum 15000 ms). There is no default wait here, so omit `wait_timeout_ms` for an immediate return.                                                                                                                                                           |

**Resource id formats.** Entry point tests use a different shape from analyses, so match carefully:

```
workspace/<ws>/app/<app>/endpoint/<METHOD>/<path>/endpointtest
workspace/<ws>/app/<app>/endpoint/FLOW/<flow_id>/endpointtest
workspace/<ws>/module/<app>/securityanalysis
workspace/<ws>/module/<app>/externalservicesanalysis
workspace/<ws>/module/<app>/buildanalysis
<any of the above>/feedback        ← feedback subresource
```

Note `app/` in the entry point test form versus `module/` in the others. Each entry point test component is one percent-escaped segment, so the path `/api/users` appears as `%2Fapi%2Fusers`. Quote ids from tool responses verbatim.

#### Job lifecycle

| Tool / Endpoint | Description                                                                                                                                                                                                                                                               |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kerno_job`     | Wait for or snapshot an async job. `wait=true` (default) blocks for at most 15 seconds, then returns the current snapshot with `wait_timed_out` if the job is still going. `wait=false` returns one immediate snapshot. Pass the `job_id` returned by the launching tool. |
| `kerno_cancel`  | Cancel in-flight async work by `job_id`, by a batch member's `resource_id`, or by `batch_id`. Fire-and-forget. The next `kerno_job` call returns `status: cancelled`.                                                                                                     |
| `GET /mcp/jobs` | HTTP endpoint. List active and recently-completed jobs as `McpJobSnapshot` objects. Supports `?since=<ISO-8601>` and `?status=<value>` (repeatable) filters.                                                                                                              |

### Job snapshots

`GET /mcp/jobs` returns an array of `McpJobSnapshot`. Use it for lightweight HTTP polling, or to retrieve a completed job's result after the MCP session has ended.

```json
{
  "jobId": "abc123",
  "kind": "endpoint_test",
  "app": "backend-api",
  "status": "healthy",
  "startedAt": "2026-07-25T10:00:00Z",
  "lastActivityAt": "2026-07-25T10:02:15Z",
  "completedAt": "2026-07-25T10:02:15Z",
  "activityLogTail": ["...last 20 lines of the activity log..."],
  "scope": {
    "endpoints": [
      { "app": "backend-api", "method": "GET", "path": "/api/items" }
    ]
  },
  "terminalPayload": { }
}
```

| Field             | Type               | Notes                                                                                                                                                          |
| ----------------- | ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `jobId`           | string             | Same ID returned by the launching MCP tool.                                                                                                                    |
| `kind`            | string             | `validate`, `endpoint_test`, `endpoint_test_many`, `healthcheck_author`, `security_analysis`, `build_analysis`, `external_services_analysis` or `calibration`. |
| `app`             | string             | Application the job targets.                                                                                                                                   |
| `status`          | string             | `running`, `needs_user_feedback` (parked at a gate, resumes once answered), `healthy`, `failed`, or `cancelled`.                                               |
| `startedAt`       | ISO-8601 timestamp |                                                                                                                                                                |
| `lastActivityAt`  | ISO-8601 timestamp | Updated on each activity log write and on completion.                                                                                                          |
| `completedAt`     | ISO-8601 timestamp | Null while running.                                                                                                                                            |
| `activityLogTail` | string\[]          | Last 20 lines of the activity log. `kerno_job` returns the full log in its `activity_log` field.                                                               |
| `scope`           | object (nullable)  | Resolves once the job starts processing. Lists the entry points targeted by this job.                                                                          |
| `terminalPayload` | object (nullable)  | Present on completion for `endpoint_test` and `validate` jobs. Holds the run's result (summary, entry point and scenario results), as `kerno_job` returns it.  |

**Query params:**

* `?since=<ISO-8601>` returns only jobs where `lastActivityAt >= since`.
* `?status=<value>` filters by status. Repeatable, as in `?status=running&status=failed`.

The last 50 finished jobs are kept. Snapshots are read-only, so use `kerno_cancel` to cancel a running job.

### MCP resources

The Kerno agent publishes MCP resources alongside its tools:

```
kerno://workspaces/{workspaceId}/scenarios
kerno://workspaces/{workspaceId}/scenarios/{+path}
kerno://workspaces/{workspaceId}/preconditions
kerno://workspaces/{workspaceId}/preconditions/{+path}
ui://kerno/graph
```

Scenario and precondition resources are scoped per workspace. The list templates return what a workspace holds, and the content templates return one file. `ui://kerno/graph` is the MCP App view that `kerno_graph` opens in hosts that support MCP Apps.

### Skills

Kerno ships its guidance for your agent as skills. Kerno installs them into your repository under `.agents/skills/`, and the MCP server also serves them through `skills/list` and `skills/get` for clients that support MCP skills. Your agent loads each one when the task calls for it.

| Skill                              | What it covers                                                                                                               |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `kerno-endpoint-test-intent`       | Choosing `effort`, `box_testing_strategy` and `tags` before the first entry point test, and checking your saved preferences. |
| `kerno-rules-template`             | A copy-paste preference table for your own rules file.                                                                       |
| `kerno-env-visibility`             | What Kerno can see in your running environment, and where saved settings and secrets go.                                     |
| `kerno-calibrate`                  | The once-per-repo calibration that runs after your application is connected.                                                 |
| `kerno-choosing-a-user-flow`       | Mapping your repository's user flows and choosing one that can run.                                                          |
| `kerno-scenario-philosophy`        | Why Kerno tests behaviour that a framework or library already enforces.                                                      |
| `kerno-presenting-results`         | Presenting a test plan awaiting approval, and run results.                                                                   |
| `kerno-presenting-diffs`           | Presenting a detected diff, and deciding with you whether to fix it or accept it.                                            |
| `kerno-presenting-coverage`        | Presenting entry point coverage from `kerno_list_endpoints`.                                                                 |
| `kerno-reviewing-lessons`          | Reviewing the lessons Kerno has learned before the next run.                                                                 |
| `kerno-baseline-coverage-at-scale` | Running baselines across a whole repository with `kerno_endpoint_test_many`.                                                 |
| `kerno-usage`                      | Writing a short Kerno usage note into your agent's own instruction files.                                                    |

### Common patterns

A typical first run looks like this.

1. `kerno_healthcheck` to confirm Docker, git, and auth.
2. `kerno_get_applications` to discover applications and pick one.
3. Start your application yourself, using its own dev flow.
4. `kerno_save_config` with an `applications` entry carrying `target_environment`, `sut_url`, and the variable names of any datastore credentials. Their secret values go in `.kerno/.env`.
5. `kerno_environment_setup`, then `kerno_environment_status` until `ready_for_endpoint_test` is true.
6. Run the `kerno-calibrate` skill, once per repository. Skip it when `.kerno/config.yaml` already holds a `calibration` record.
7. Map your user flows with the `kerno-choosing-a-user-flow` skill, then save the ones you confirm with `kerno_save_flows`.
8. `kerno_endpoint_test` with `type: generate` for the entry point you care about, by method and path. `kerno_list_endpoints` lets you browse what analysis found and which entry points already have tests.
9. Block on the launch response's `monitoring_command` until it exits at the plan-approval gate. `kerno_job` reports the same gate as `needs_user_feedback`.
10. Read the plan with `kerno_feedback_pending`, then approve it via `kerno_feedback_answer`.

After a code change, call `kerno_sync_workspace`, then `kerno_validate`. Pass the entry points that `kerno_list_endpoints` returns with scope `changed`, and its `impactedFlows` as `flows`, or leave out both to validate every known entry point. If an entry point changed on purpose and its tests should follow, run `kerno_endpoint_test` with `type: update`.

### Notes for agent integrators

* Use absolute workspace paths when tools request `workspace_path`. It is the selector for which workspace the call targets, so pass the one you mean on every call. The agent accepts its startup workspace and any directory nested under it, and holds several at once.
* Do not poll `kerno_job` in a tight loop. Let `wait=true` do the waiting, up to 15 seconds per call, block on `kerno_get_state` or `kerno_await_state`, or use `GET /mcp/jobs`.
* **Feedback gates show on the job.** A gated entry point test reports `needs_user_feedback` on `kerno_job`, and its state resource carries `open_feedback`.
* Do not manually edit scenario files under `.kerno/scenarios/`. To change them, run `kerno_endpoint_test` with `type: update` or `type: generate`.
* A second launch of the same entry point test with the same settings is refused while the first is running, and a batch rejects any member whose entry point is already under test.
* After making code changes, call `kerno_sync_workspace` before `kerno_list_endpoints` or `kerno_endpoint_test`.
* While an entry point test runs, block on its `monitoring_command`. The command needs `curl` and `jq`. Without them, poll `kerno_feedback_pending` and `kerno_get_state` yourself.
* For long-poll calls, `timed_out=true` is a normal outcome, so re-issue immediately with the same parameters.


# Testing modes

Every dial on an entry point test: what to generate, how hard to try, and how much access Kerno gets.

These dials set the ratio of test quality to speed you want, and how much access Kerno gets while doing it.

`kerno_endpoint_test` takes one required dial and several optional ones. The defaults are deliberately thorough, so you only need to reach for these when you want something faster, narrower, or more cautious.

{% hint style="info" %}
**Kerno stores none of these preferences.** If you settle on a set you like, your agent writes them into your own rules file so they are diffable, reviewable in a pull request, and travel with the repo. See [Saving your preferences](#saving-your-preferences).
{% endhint %}

### `type`, what the run does

Required. No default.

| Value      | What it does                                                                                                                             |
| ---------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| `generate` | Analyzes the entry point, plans scenarios, **pauses for your approval**, then implements and runs them. Use this the first time.         |
| `update`   | Conservatively repairs existing scenarios after an intentional change. Keeps what still applies and edits only what the change requires. |

`update` needs scenarios to exist. If none do, the run ends immediately with `needs_generate` and a message telling you to generate first.

To re-run the scenarios already on disk after a code change, use the separate `kerno_validate` tool. It runs existing scenarios only, with no planning and no implementation. See [Kerno MCP](/docs/references/kerno-mcp).

**Choosing between `update` and `generate`.** Use `update` when your entry point changed on purpose and the tests should follow: it defaults every existing scenario to "keep", makes surgical edits, and never deletes the happy path. Use `generate` when you want to start the entry point's coverage over.

`generate` always pauses for plan approval. `update` only pauses if you set `interactive: true`.

### `effort`, how hard Kerno tries

Optional. Defaults to **`high`**. An entry point your team marked critical in the Portal always runs at `high`, whatever you pass.

| Value    | Runs per scenario                 | Critique and repair | What you get                                                                                                                                            |
| -------- | --------------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `low`    | 1                                 | no                  | The fastest, cheapest signal. The scenario is written and proven to pass once.                                                                          |
| `medium` | 2 consecutive, both must be clean | no                  | Proof the scenario is genuinely re-runnable, plus a second sample of randomized test data.                                                              |
| `high`   | 2 consecutive                     | yes                 | Everything `medium` gives, plus an implement-critique-repair loop that rejects weakened, tautological, or swallowed assertions and incomplete coverage. |

The second run is not noise suppression. If a scenario passes once and fails the second time, that is a real finding: either its own setup and cleanup are not isolated, or the fresh random data hit an edge case the first sample missed. Both are worth knowing.

{% hint style="info" %}
`kerno_validate` re-runs the scenarios already on disk and takes no `effort` argument.
{% endhint %}

### `box_testing_strategy`, how much access Kerno gets

Optional. Defaults to **`white_box`**.

|                                              | `black_box`                | `white_box`                                                               |
| -------------------------------------------- | -------------------------- | ------------------------------------------------------------------------- |
| Datastore drivers available to scenario code | no                         | yes                                                                       |
| Database schema in the planning prompt       | suppressed                 | included                                                                  |
| Assertions about stored state                | not planned                | planned where useful                                                      |
| How preconditions and results are handled    | through your HTTP API only | through the API, falling back to the datastore where the API cannot do it |

**Black box** means Kerno only talks to your application over HTTP. Every precondition is established and every result verified through your own entry points.

**White box** means Kerno may additionally read and write your datastore directly, for the cases your API cannot express: seeding a record your API has no way to create, or checking that a stored password is not the plaintext that was sent. This is what the datastore blocks in `kerno_save_config` enable. See [Environment Setup](/docs/core-concepts/environment-setup) for wiring them up.

Two things that surprise people:

* **Your datastore credentials still reach the sandbox under `black_box`.** The dial works by controlling planning and access. Drivers are withheld, the schema is suppressed, and preconditions and results run over HTTP. The environment still receives your credentials.
* **`black_box` does not hide your source code** from Kerno's planner. It constrains how scenarios interact with your running system, not what Kerno reads while writing them.

Configuring datastore access does not lock you into white box. Any individual run can opt out by passing `box_testing_strategy: black_box`.

{% hint style="warning" %}
**White box needs a schema it can derive from your source.** To read or write your datastore directly, Kerno first works out its shape from your source code, for example schema files, migrations, or ORM models. If it cannot derive one, any scenario that needs direct datastore access is marked blocked rather than run, and Kerno may ask you once to point it at a schema. Scenarios that only exercise HTTP are unaffected.
{% endhint %}

### `tags`, what kind of coverage

Optional. Defaults to `validation`.

| `tags`                      | Happy path | Functional coverage | Security scenarios |
| --------------------------- | ---------- | ------------------- | ------------------ |
| omitted or `["validation"]` | yes        | yes                 | no                 |
| `["security"]`              | yes        | **no**              | yes                |
| `["validation","security"]` | yes        | yes                 | yes                |

Note the middle row: `security` on its own **narrows** the run rather than adding to it. If you want your normal functional coverage plus security scenarios, pass both tags.

The happy path is always planned, whatever you pass.

See [Security testing](/docs/references/security-testing) for what the `security` tag actually produces.

### `test_generation_context`, guidance for this entry point

Optional free text describing what Kerno should know or do differently for this entry point.

Its meaning depends on `type`:

* **`generate`.** Complete guidance for planning. Passing a **different** value than last time discards the scenarios on disk and re-plans from scratch, so resend your previous guidance in full when you are adding to it.
* **`update`.** Change intent: what changed and how the tests should adapt. It does not wipe existing scenarios.

**Relationship to workspace config.** Durable rules belong in your Kerno config under `test-generation`, where they are version-controlled and apply automatically. See [Custom rules](/docs/core-concepts/custom-rules). The two are combined for each run, with this per-call argument taking precedence where they conflict.

{% hint style="warning" %}
Only the per-call argument triggers a re-plan. Editing `test-generation` rules in `.kerno/config.yaml` changes the guidance Kerno receives on the next `generate`, but existing scenarios are kept rather than re-planned. To force config changes through, either pass a changed `test_generation_context` or remove the entry point's scenarios first.
{% endhint %}

### `interactive`, step through the run

Optional, defaults to `false`. When true, Kerno pauses before implementing each scenario to ask whether to proceed, skip, or take feedback, and pauses again if automatic retries are exhausted. For `type: update`, it also restores the plan-approval gate.

These pauses are gates that wait for your answer. `kerno_job` reports a paused run as `needs_user_feedback` with the open question, and `kerno_feedback_pending` lists it too.

### `scenario_ids`, target specific scenarios

Optional. When provided, only the named scenarios are implemented or run; the rest are skipped. Omit it to target every scenario.

Note that this filters implementation and execution but not planning: a `generate` run still plans the full set. If none of the ids match, the call fails rather than quietly doing nothing.

### Saving your preferences

Kerno is stateless about your habits. It stores no testing preferences, and it does not learn them silently over time.

Instead, every entry point test response echoes a `resolved_intent` object showing what was used and which values had to be defaulted:

```json
"resolved_intent": {
  "effort": "high",
  "box_testing_strategy": "white_box",
  "tags": ["validation"],
  "defaulted": ["effort", "box_testing_strategy", "tags"]
}
```

When anything was defaulted, it also carries a `save_to_rules_hint`. When a critical mark raised the effort, it carries `effort_forced_by: "critical_endpoint"`. `kerno_endpoint_test_many` reports only `effort_forced_by`, for each entry point.

If you have a standing preference, your coding agent writes it into your own rules file, `CLAUDE.md`, `.cursor/rules`, or `AGENTS.md`, as a pattern-matched table. That keeps it diffable, reviewable in a pull request, and travelling with the repository.

Ask your agent to use the `kerno-rules-template` skill to get the table format. The rule is that the most specific entry point pattern wins, and you should always keep one `*` fallback row describing what a bare call should resolve to for your codebase.


# Security testing

What the security tag produces, how Kerno picks the vulnerabilities worth testing, and how to review what it flags.

Security tests baseline your entry point's security posture. Kerno records how the entry point stands up to a specific class of attack today, so if a later change opens that vulnerability, the test starts failing and the regression is caught in the same loop as any other behaviour change.

### Turning it on

Security coverage is controlled by the `tags` argument on `kerno_endpoint_test`. It defaults to `validation`.

| `tags`                      | Happy path | Functional coverage | Security scenarios |
| --------------------------- | ---------- | ------------------- | ------------------ |
| omitted or `["validation"]` | yes        | yes                 | no                 |
| `["security"]`              | yes        | **no**              | yes                |
| `["validation","security"]` | yes        | yes                 | yes                |

Note the middle row. `security` on its own **narrows** the run rather than adding to it, so pass both tags when you want your normal functional coverage alongside security scenarios. The happy path is always planned, whatever you pass.

Ask for it through your agent:

```
Use Kerno to generate tests for POST /api/users with validation and security coverage.
```

### What the security tag changes

It changes two stages of the run.

**Planning** assesses which **OWASP API and Web Top 10** categories plausibly apply to the entry point, given its inputs, its authentication model, and what the surrounding code does, then plans tests only for the categories that fit. Common ones include broken object-level and function-level authorization, injection, excessive data exposure, mass assignment, and server-side request forgery.

**Implementation** frames the baseline assertions as vulnerability checks. A passing scenario means the entry point is not vulnerable to that case today, and a later baseline diff explains in plain English what the vulnerability is.

That framing is why a security scenario is worth keeping even while it passes. The value is in the day it stops passing.

### Reading the results

Security scenarios report the same four verdicts as any other scenario, described in [Baseline tests](/docs/core-concepts/scenarios-and-baselines#reading-test-results). A failing security scenario means the entry point's security behaviour moved, and it is for you to decide whether the new behaviour is intended.

### Potential bugs

A scenario that passes can still carry a **potential bug**. A run saw the entry point do something the plan did not expect, and Kerno wrote the test to document that real behaviour. Kerno records one only when a run actually observed it. The flag names what deviated and what a future change would signify. When it names a suspected code location, the location counts as verified only if Kerno read that code during the run.

If you review one and decide the behaviour is known or intended, tell your agent to ignore it:

```
The role field Kerno flagged on GET /users/:id is intentional, we keep returning it
for backwards compatibility. Have Kerno ignore that potential bug.
```

Kerno records the decision for that entry point, alongside the code revision you made it against, and later run reports leave it out. The test file keeps its flag until the entry point's tests are next updated, and a full cache clear discards the decision.

{% hint style="warning" %}
There is currently no way to reverse an ignore decision through Kerno, so read the flag carefully before you dismiss it.
{% endhint %}

### Related

* [Testing modes](/docs/references/testing-modes) for every other dial on a run.
* [Baseline tests](/docs/core-concepts/scenarios-and-baselines) for what a baseline records.
* [Supported technologies](/docs/references/supported-technologies) for the authentication mechanisms Kerno handles.


# Anatomy of a Kerno test

What Kerno writes into your repository for each entry point, which file to read, and which one to leave alone.

Kerno writes its tests into your repository so they live with your code and travel through review like anything else. Each entry point gets its own folder.

The important thing to know first: **every scenario exists as two files.** The `.scenario.md` is written for you, and the `.scenario.ts` is written for the agent that runs it. When you want to know what a test does, read the Markdown.

### What is in an entry point's folder

For `POST /api/auth/login`, Kerno creates `.kerno/scenarios/endpoints/POST/api/auth/login/`:

| File                      | What it is                                                                                               | Read it? |
| ------------------------- | -------------------------------------------------------------------------------------------------------- | -------- |
| `<name>.scenario.md`      | The scenario in plain English. One per scenario                                                          | **Yes**  |
| `<name>.scenario.ts`      | The same scenario as executable TypeScript, which is what actually runs                                  | Rarely   |
| `preconditions.ts`        | Checks that run before the entry point's scenarios, confirming the entry point is reachable and testable | Rarely   |
| `preconditions.md`        | The same checks in plain English                                                                         | Rarely   |
| `plan.json`               | The plan the Markdown is rendered from                                                                   | No       |
| `report.json`             | Results from the most recent run                                                                         | No       |
| `<name>.scenario.blocked` | Marks a scenario that could not run                                                                      | No       |

Everything here is committed alongside your code.

### The Markdown, which is the one to read

Each `.scenario.md` opens with an HTML comment marking it as generated, then follows the same section order every time. This is a real one, with the credential masked:

```markdown
## Successful login with valid credentials returns tokens and user data
**Kind:** HAPPY_PATH | **Actor:** unauthenticated visitor | **Scope:** POST /api/auth/login

### Preconditions
- The database contains an active user with email `admin@example.com` and password hash matching `<password>`

### Trigger
POST /api/auth/login with JSON body `{ "email": "admin@example.com", "password": "<password>" }`

### Main success scenario
1. The endpoint validates the email format and password presence
2. The endpoint finds the active user by email
3. The endpoint verifies the plaintext password against the stored hash
4. The endpoint generates a JWT access token and a refresh token
5. The endpoint updates the user's `refreshToken` column and `lastLoginAt` timestamp
6. The endpoint responds 200 with JSON containing `user`, `token`, and `refresh_token`

### Postconditions
- Response status is `200`
- Response body `user.email` equals `"admin@example.com"`
- Response body `user.id` is a UUID string
- Response body `token` is a non-empty string
- The `token` value decodes as a valid JWT (does not throw on decode)
- A `user` table row with email `admin@example.com` exists with a non-null `refreshToken` column
- The same row's `lastLoginAt` column is set to a timestamp within the last 10 seconds

### Notes
The endpoint is throttled at 5 requests per 60 seconds per the `@Throttle` decorator. The
stored hash may be bcrypt or scrypt; the service auto-migrates to the current algorithm on
successful login.
```

Reading top to bottom you get what is worth knowing about any scenario. Sections marked optional appear only when the plan has them.

| Section                   | What it tells you                                                                                        |
| ------------------------- | -------------------------------------------------------------------------------------------------------- |
| **Kind / Actor / Scope**  | What sort of case this is, who is making the request, and which entry point it hits                      |
| **Preconditions**         | The state that has to exist before the request is sent                                                   |
| **Trigger**               | The exact request                                                                                        |
| **Main success scenario** | What your code is expected to do, step by step                                                           |
| **Postconditions**        | Every assertion, in plain English. This is the baseline                                                  |
| **Minimal guarantees**    | Optional. What must hold even when the scenario fails                                                    |
| **Cleanup**               | Optional. The teardown that runs after the assertions                                                    |
| **Notes**                 | Optional. What Kerno noticed and deliberately decided about, including what it chose to leave unasserted |

The **Postconditions** section is the one to read closely. It is the baseline in words, so when a validation run reports a diff, the assertion that moved is written there.

The **Notes** section is worth a look when a scenario surprises you. Kerno records the judgement calls it made, such as an optional field it saw and chose to leave out of the assertions.

### The TypeScript, which the agent maintains

The `.scenario.ts` is what the sandbox executes. It is a Vitest file that talks to your running application over HTTP and, in white box mode, to your datastores directly.

```typescript
import { ctx } from '@kerno/ts-sandbox'
import { expect } from 'vitest'
import pg from 'pg'

export const meta = {
  description: `POST /api/auth/login with the seeded credential returns 200 with a user
object, a JWT token, and a refresh_token; the user row's refreshToken and lastLoginAt are updated.`,
  path: 'POST /api/auth/login',
}

export default function () {
  const credentials = { email: 'admin@example.com', password: '<password>' }
  let loginResponse: { status: number; body: any }

  ctx.act('POST /api/auth/login, then read back the user row', async () => {
    const res = await fetch(`${process.env.SUT_BASE_URL}/api/auth/login`, {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify(credentials),
    })
    loginResponse = { status: res.status, body: JSON.parse(await res.text()) }
    // ... reads the user row directly from Postgres to verify the write
  })

  ctx.assert('Response and persisted row match the login contract', () => {
    expect.soft(loginResponse.status, 'POST /api/auth/login responds 200').toBe(200)
    expect.soft(loginResponse.body, 'returns exactly user/token/refresh_token').toEqual({
      user: expect.anything(),
      token: expect.anything(),
      refresh_token: expect.anything(),
    })
    // ... one soft assertion per postcondition
  })

  ctx.cleanUp('No persistent state created', async () => {
    // login only updated the existing seeded user
  })
}
```

A scenario runs in up to four phases, and each one is a `ctx` call:

| Phase         | What it does                                 |
| ------------- | -------------------------------------------- |
| `ctx.arrange` | Sets up the state the Preconditions describe |
| `ctx.act`     | Sends the request                            |
| `ctx.assert`  | Checks the Postconditions                    |
| `ctx.cleanUp` | Removes whatever the scenario created        |

Two details explain why the generated code looks the way it does.

**Assertions are soft.** `expect.soft` lets a run finish and report every mismatch at once, so one validation tells you everything that moved instead of stopping at the first difference.

**Every assertion carries a message.** The second argument to each `expect` is the plain-English form of that postcondition, which is what you see in a run report when it fails.

### Flow tests

A [user flow](/docs/core-concepts/user-flows) gets one test that walks the journey in order. It lives under `.kerno/scenarios/flows/<flow>/`, beside a `flow.json` that records the flow's description, its steps and the outcome it should reach. In the TypeScript, each step of the journey runs inside its own `ctx.step(...)`, and the test checks the outcome at the end. When a journey breaks, the result names the step where it failed.

### Preconditions

`preconditions.ts` runs before the entry point's scenarios and checks, read-only, that testing the entry point is possible at all. What it checks depends on the entry point.

* **HTTP route.** The application URL is set, the route is mounted and answering, and any credential the scenarios need can be obtained.
* **MCP tool.** One read-only call to the tool answers.
* **Background consumer.** Only the environment and seed data below, since a consumer is triggered through its broker and nothing signs in.
* **User flow.** The service answers and the flow's actor can sign in.

Each also checks that the environment variables the run needs are set, and routes, tools and consumers check that their seed data exists. Reaching your databases and other dependencies is the environment healthcheck's job.

When a precondition fails, the run stops before any scenario runs and ends with an error naming the reason, which tells you nothing was tested. In an interactive run, Kerno first asks whether to retry with your guidance or abort. Validation runs reuse the tests on disk and skip preconditions.

### Changing a test

Ask Kerno rather than editing the files.

```
Use Kerno to update the tests for POST /api/auth/login.
The response no longer includes the role field.
```

Kerno re-plans, re-renders the Markdown, and rewrites the TypeScript together, so the two halves stay in step. An edit you make by hand is overwritten the next time that entry point is generated or updated.

### Related

* [Baseline tests](/docs/core-concepts/scenarios-and-baselines) for what a baseline records and how results are reported.
* [Testing modes](/docs/references/testing-modes) for the dials that change what gets generated.
* [Ports and file paths](/docs/references/ports-and-file-paths) for where everything lives on disk.


# Scopes

A scope tells Kerno which entry points to list.

A scope narrows Kerno's view of your application to a subset of entry points. It is a required argument on `kerno_list_endpoints`.

### The four scopes

| Scope                     | Meaning                                                |
| ------------------------- | ------------------------------------------------------ |
| `all`                     | Every entry point in the selected application          |
| `changed`                 | Only entry points affected by your current git changes |
| `file:path/to/handler.ts` | Entry points defined in that source file               |
| `endpoint:METHOD /path`   | A single route, identified by HTTP method and path     |

### When to use each

* **`all`.** Surveying an application, or checking overall coverage.
* **`changed`.** The everyday one. Kerno compares your working tree against `HEAD`, including both staged and unstaged changes. A new file counts once it is staged, and a change stops counting once you commit it. Kerno resolves which entry points your edits reach. It follows the call graph rather than just the file you edited, so a change to a shared helper surfaces every entry point downstream of it. Touched scenario files count too, not only handlers. It also reports the [user flows](/docs/core-concepts/user-flows) that cross those entry points, so they can be re-run too.
* **`file:`.** Focused changes to one handler file.
* **`endpoint:`.** A single route you already have in mind.

### Syntax

`file:` paths are interpreted relative to your workspace root.

`endpoint:` expects exactly one space between the method and the path, as in `endpoint:POST /api/orders`. The method is uppercased for you and the path normalized, but both must be present.

### Where scopes apply

Scopes are used by **`kerno_list_endpoints`**, where the argument is required.

{% hint style="info" %}
`kerno_endpoint_test` does not take a scope. It targets one entry point at a time, identified by an explicit `endpoint_method` and `endpoint_path`. `kerno_endpoint_test_many` takes an explicit list of entry points, also without a scope. The usual pattern is to list entry points with a scope to see what needs attention, then run a test against the ones you picked.
{% endhint %}

### How scopes work behind the scenes

Kerno's dependency graph maps every entry point to the code that produces it. When you pick a scope, Kerno resolves it to a concrete list of entry points and works from that list.

See [Change Validation](/docs/core-concepts/change-validation) for how the `changed` scope fits into the everyday loop.


# Ports and file paths

This is the reference for the network ports and filesystem paths Kerno uses on your machine.

### Ports

| Service                            | Default port | Override             |
| ---------------------------------- | ------------ | -------------------- |
| Agent REST API (used by the CLI)   | `8085`       | `AGENT_PORT` env var |
| MCP server (used by coding agents) | `8086`       | `MCP_PORT` env var   |

Both ports listen on localhost only and are not reachable from other machines on your network.

If the default port is taken, Kerno auto-allocates a free one. The actual port is printed by `kerno init`.

{% hint style="warning" %}
Treat the defaults above as documentation, not as addresses to hard-code. Kerno reuses the default ports across restarts when it can, and falls back to a free port only when a default is unavailable, so the actual port can differ. Always copy the URL from your terminal rather than reusing one from an old config file.
{% endhint %}

### File paths

Kerno uses these locations on your filesystem.

#### `~/.kerno/`

Installed agent binaries, runtime state, and per-workspace data.

| Path                                            | Purpose                                                                 |
| ----------------------------------------------- | ----------------------------------------------------------------------- |
| `~/.kerno/assets/agent/{version}/aicore-agent/` | Installed agent distribution                                            |
| `~/.kerno/assets/runtime/{version}/custom-jre/` | Custom JRE bundled with the agent                                       |
| `~/.kerno/assets/scip/`                         | Code indexers and the Node runtime they need                            |
| `~/.kerno/assets/index/`                        | Type packages downloaded for code indexing                              |
| `~/.kerno/agent/`                               | Agent working directory, cleared on each agent start                    |
| `~/.kerno/workspaces/<workspace>/`              | Per-workspace snapshots, indexed data, prompt cache, and logs           |
| `~/.kerno/agent.stdout.log`                     | Agent startup output, the first place to look when the agent won't boot |
| `~/.kerno/state.json`                           | CLI state, including your selected organization                         |
| `~/.kerno/cli-update-check.json`                | When the CLI last checked npm for a newer version                       |
| `~/.kerno/cli.installed`                        | Marks that the CLI has run on this machine                              |

**Runtime state files.** While the agent is running, `~/.kerno/` also holds `agent.pid`, `agent.port`, `mcp.port`, `agent.workspace` and `agent.version`.

{% hint style="info" %}
These exist only while the agent is up. If you go looking for `mcp.port` and it is not there, the agent is not running. That is expected.
{% endhint %}

**Credentials.** Your login tokens are held in the OS keychain (macOS Keychain, the Linux Secret Service via `secret-tool`, or Windows Credential Manager). Only when no keychain is available do they fall back to `~/.kerno/secrets.json`, readable only by you.

#### `.kerno/` (in your repository)

Kerno creates this in the workspace your agent works in, normally your repository root. In a monorepo, scenarios live under each app's own directory. A nested directory your agent uses as its workspace gets its own `.kerno/config.yaml`, and settings saved there live there.

Kerno auto-writes a `.kerno/.gitignore` containing one entry, `.env`, so your secret values stay out of your history. The rest of `.kerno/`, including `config.yaml`, is intended to be committed alongside your code.

**At the workspace root**

| Path                                 | Purpose                                                                                                                                        | Committed   |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------- | ----------- |
| `config.yaml`                        | How Kerno reaches your application, and the name of the environment variable holding each dependency's credentials. `config.yml` also works    | yes         |
| `.env`                               | The real values for the variables `config.yaml` names, one `VAR=value` per line                                                                | no          |
| `.gitignore`                         | Written by Kerno, containing `.env`                                                                                                            | yes         |
| `healthcheck.<local/remote>.ts`      | The environment healthcheck [calibration](/docs/core-concepts/environment-setup#calibration) generates from your infrastructure. Yours to edit | yes         |
| `CHANGES_DETECTED.md`                | Entry points impacted by your current changes, written by a `changed`-scope run                                                                | yes         |
| `memory/`                            | What Kerno has learned about your codebase, one Markdown file per record plus a README, for review in pull requests                            | yes         |
| `config.yaml.unreadable-<timestamp>` | A `config.yaml` Kerno could not parse, moved aside before it writes a new one                                                                  | your choice |
| `monitor-endpoint-test-job.sh`       | A helper script your agent uses, kept out of git                                                                                               | no          |
| `settings.yaml`                      | Optional. Each `key: value` under its `agent:` section is passed to the agent as an environment variable when the CLI starts it                | your choice |

{% hint style="warning" %}
**`config.yaml` is committed.** Put real secret values in `.kerno/.env`, which Kerno gitignores for you. `config.yaml` records only the *name* of the environment variable each secret uses, so it is safe to review in a pull request alongside your code.

Kerno merges `.env` into the sandbox environment when a test runs, and a value there overrides one of the same name in `config.yaml`.
{% endhint %}

**What `config.yaml` holds**

One file in the workspace covers every application Kerno found there, keyed by application id under `applications`. Each entry carries that application's own URL and its own dependency access, which is how a monorepo runs several services at once.

```yaml
# .kerno/config.yaml
applications:
  payments-api:
    target-environment: "local"
    sut:
      input-url: "http://localhost:8081"
    infrastructure:
      postgres:
        access: "direct"
        env:
          DATABASE_URL: "<set in .kerno/.env>"

  orders-api:
    target-environment: "local"
    sut:
      input-url: "http://localhost:8082"
    infrastructure:
      postgres:
        access: "direct"
        env:
          DATABASE_URL: "<set in .kerno/.env>"
      redis:
        access: "none"
        reason: "Shared with staging, so Kerno leaves it alone"
```

| Key                     | What it sets                                                                                                                                                                                                                                                                                                                                                                |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `target-environment`    | `local` when the application runs on your machine, `remote` when it runs somewhere you reach                                                                                                                                                                                                                                                                                |
| `sut.input-url`         | Where that application is listening                                                                                                                                                                                                                                                                                                                                         |
| `infrastructure.<kind>` | One entry per [component](/docs/references/supported-technologies#supported-external-dependencies) the application talks to. Its key is one of `postgres`, `mariadb`, `mysql`, `mongodb`, `redis`, `kafka`, `rabbitmq`, `azure-servicebus`, `aws`, `aws-sqs`, `dynamodb`, `clickhouse`, `azurite`, `s3`, `zitadel`, `cognito` or `bifrost`, and Kerno ignores any other key |
| `<kind>`                | An older form some saves still write directly under the application, such as `postgres:` followed by its variables. It reads as `access: direct`, and an `infrastructure` entry for the same kind takes precedence                                                                                                                                                          |
| `http-services`         | Other HTTP services the application calls, each with a `base-url` and an `openapi-spec`. Also allowed at the top level                                                                                                                                                                                                                                                      |
| `test-generation`       | [Custom rules](/docs/core-concepts/custom-rules) for that application                                                                                                                                                                                                                                                                                                       |

A key Kerno does not define is kept as written and has no effect, so a typo in a hand-edited key changes nothing. When your agent saves settings, Kerno names any such keys.

At the top level, beside `applications`, a `calibration` block records the outcome of [calibration](/docs/core-concepts/environment-setup#calibration), either `calibrated` or `declined`, with the date and a short note. Calibration writes it, and your agent reads it to know the repository is already calibrated.

Each `infrastructure` entry sets `access` to one of two values. Use `access: direct` with an `env` block naming the variables Kerno needs, and it will connect to that component. Use `access: none` on its own, with an optional `reason`, to record that the component exists and have Kerno leave it alone.

{% hint style="info" %}
The `env` block names the variables. Put the real values in `.kerno/.env`, which Kerno gitignores for you and merges in when a test runs, overriding anything of the same name here.
{% endhint %}

**Per application**

| Path                                       | Purpose                                                                                                                    | Committed |
| ------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------- | --------- |
| `scenarios/endpoints/<METHOD>/<path>/`     | One folder per entry point, holding its scenarios. See [Anatomy of a Kerno test](/docs/references/anatomy-of-a-kerno-test) | yes       |
| `scenarios/endpoints/.../*.scenario.md`    | Each scenario in plain English. The one to read                                                                            | yes       |
| `scenarios/endpoints/.../*.scenario.ts`    | Each scenario as executable TypeScript                                                                                     | yes       |
| `scenarios/endpoints/.../preconditions.ts` | Checks that run before the entry point's scenarios                                                                         | yes       |
| `scenarios/endpoints/.../plan.json`        | The plan the Markdown is rendered from                                                                                     | yes       |
| `scenarios/endpoints/.../report.json`      | Results from the most recent run                                                                                           | yes       |
| `scenarios/flows/<flow>/`                  | One folder per [user flow](/docs/core-concepts/user-flows), holding the flow's record and its test                         | yes       |
| `analysis/`                                | Per-entry-point security recipes as JSON, reused as context across runs                                                    | yes       |
| `criticality.json`                         | The entry points your team marked critical in the Portal, synced into the repository                                       | yes       |

#### `/tmp/kerno/`

Temporary working directories.

| Path                                | Purpose                                                    |
| ----------------------------------- | ---------------------------------------------------------- |
| `/tmp/kerno/scenario-workspace/...` | Working directory for scenario planning and implementation |

#### `$TMPDIR`

| Path                                  | Purpose                                                         |
| ------------------------------------- | --------------------------------------------------------------- |
| `$TMPDIR/kerno/sandbox/<project>`     | The scenario sandbox's own working files                        |
| `$TMPDIR/code-index-<repo>-<commit>/` | Per-commit checkouts used to build the code index               |
| `$TMPDIR/kerno-tool-results/`         | Large tool results set aside during a run, readable only by you |

#### Agent skills (in your repository)

Kerno installs its agent skills at `.agents/skills/kerno-*`, with links to them at `.claude/skills/kerno-*`. It keeps them out of git by adding `.agents/skills/*`, the links and `.kerno/monitor-endpoint-test-job.sh` to `.git/info/exclude`. The first line covers the whole folder, so any other skill you keep in `.agents/skills/` is ignored too.

### Removing Kerno

To stop the agent and remove the installed binaries, run:

```bash
kerno uninstall
```

This stops the running agent and clears `~/.kerno/assets/`, which includes the agent distribution, the bundled JRE, and the code indexers.

Other state is preserved unless you delete it manually: `~/.kerno/workspaces/`, your repository's `.kerno/` directory, anything under `/tmp/kerno/`, and your login credentials in the OS keychain.


# Supported Technologies

Overview of Kerno's supported technologies

### **Supported Coding Agents**

Kerno works with any coding agent that supports MCP, including Claude Code, Cursor and Codex. See the [Quickstart](/docs/getting-started/quickstart) to connect yours. We're constantly adding new technologies, so if you want us to support your stack next, [let us know here](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ).

### **Supported Backend Languages**

Your scenarios are always written in TypeScript and reach your application the way its clients do, over HTTP for routes, through an MCP client for MCP tools, and through your message broker for background consumers. The language your service is written in does not affect how tests run. What it affects is how Kerno discovers your routes.

For the frameworks below, Kerno detects routes from your code. When detection finds no routes in an application, Kerno's analysis searches the code for them itself. A route that detection misses in an application where it finds others stays off the list until you ask for a test on it by method and path, and then Kerno locates its handler.

| Language                | Frameworks with route detection                                                                                                                                    |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| TypeScript / JavaScript | Express, NestJS, Fastify, Koa, Hono, Hapi, AdonisJS, Elysia, Next.js App Router and Pages Router API routes, Remix / React Router v7, tsoa, SvelteKit, Nuxt, Astro |
| Python                  | Django, Flask, FastAPI, Litestar, Sanic, Tornado, Pyramid, Bottle, Falcon, aiohttp                                                                                 |
| Java                    | Spring, Quarkus, Jersey, RESTEasy, Dropwizard, Micronaut, Vert.x Web                                                                                               |
| Kotlin                  | Ktor, Spring, Micronaut                                                                                                                                            |
| Scala                   | http4s, Akka HTTP, Pekko HTTP, Play                                                                                                                                |
| Go                      | Gin, Echo, Chi, Fiber, GoFrame, gorilla/mux, net/http                                                                                                              |
| Ruby                    | Rails, Sinatra, Grape                                                                                                                                              |
| PHP                     | Laravel, Symfony, Slim, Drupal, Utopia (Appwrite)                                                                                                                  |
| C#                      | ASP.NET, FastEndpoints                                                                                                                                             |
| Rust                    | Axum, actix-web, Rocket                                                                                                                                            |
| Swift                   | Vapor                                                                                                                                                              |

For SvelteKit, Nuxt and Astro, detection reads page and API paths from the file layout without their HTTP method, and SvelteKit `+server` endpoints go undetected. Ask for a test on one by method and path, and Kerno locates its handler.

Kerno also reads AWS deployment files. HTTP routes declared in an AWS SAM template are detected, and SQS-triggered functions in a Serverless Framework `serverless.yml` are detected as background consumers.

Don't see your framework? [Let us know](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ). Route detection is added as plugins, so new frameworks land quickly.

### **Supported External Dependencies**

These are the dependencies Kerno can connect to directly, for setting up and verifying state that your API cannot express. See [Environment Setup](/docs/core-concepts/environment-setup).

| Dependency                                         | Status    |
| -------------------------------------------------- | --------- |
| PostgreSQL                                         | Supported |
| MariaDB                                            | Supported |
| MySQL                                              | Supported |
| MongoDB                                            | Supported |
| Redis                                              | Supported |
| Kafka                                              | Supported |
| RabbitMQ                                           | Supported |
| Azure Service Bus                                  | Supported |
| Amazon SQS                                         | Supported |
| DynamoDB                                           | Supported |
| ClickHouse                                         | Supported |
| Azure Blob, Queue, and Table Storage (via Azurite) | Supported |
| S3-compatible object storage (MinIO and similar)   | Supported |
| Zitadel                                            | Supported |
| Amazon Cognito                                     | Supported |

Kerno connects to these over the network, so embedded databases such as SQLite cannot be accessed directly. Applications backed by them are still fully testable through their HTTP API.

Zitadel and Amazon Cognito are identity providers rather than stores. Kerno reads them to obtain a test credential, and seeds no state in them.

For AWS-backed dependencies, an `aws` block carries your shared credentials and region, which every AWS SDK client a scenario builds then picks up. The individual kinds, such as `aws-sqs` and `dynamodb`, declare what they each need on top of that.

### **Supported Protocols**

| Protocol                                                        | Status                                        |
| --------------------------------------------------------------- | --------------------------------------------- |
| HTTP (S)                                                        | Supported                                     |
| MCP (Streamable HTTP)                                           | Supported                                     |
| Background consumers (Celery, Graphile Worker, broker messages) | Supported                                     |
| <mark style="color:$info;">WebSocket</mark>                     | <mark style="color:$info;">Coming Soon</mark> |
| <mark style="color:$info;">gRPC</mark>                          | <mark style="color:$info;">Coming Soon</mark> |
| <mark style="color:$info;">GraphQL</mark>                       | <mark style="color:$info;">Coming Soon</mark> |

### **Supported Authentication Methods**

Kerno reads your source code to work out how each entry point authenticates, then builds a per-entry-point recipe for obtaining and presenting a credential. These are the mechanisms it handles reliably.

| Authentication Methods      | Status    |
| --------------------------- | --------- |
| Bearer Authentication (JWT) | Supported |
| Session Tokens              | Supported |
| API Key (Custom Header)     | Supported |
| HTTP Basic Authentication   | Supported |
| OAuth 2.0                   | Supported |

{% hint style="info" %}
When Kerno signs a token itself, it uses HMAC algorithms (`HS256`, `HS384`, `HS512`). If your verifier requires an asymmetric algorithm such as `RS256`, Kerno obtains a real token from your login flow instead of constructing one.

Flows that require a person to complete a hand-off with an external identity provider cannot be automated. Kerno reports those scenarios as blocked rather than faking a credential.
{% endhint %}

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we'll gladly help.
{% endhint %}


# Security & Privacy

Understand how Kerno handles your code and data.

Kerno is designed to keep your code local. This page explains what data stays on your machine, what gets sent to LLM providers, and what Kerno collects to power its services.

### Your code stays local

The Kerno agent runs on your machine. Your repository, dependency graph, test environment, scenarios, and baselines all live on your filesystem and inside your local Docker containers. Kerno never copies your code to its servers for storage.

### What leaves your machine

Kerno uses large language models to generate scenarios and reason about your code. The agent sends relevant excerpts of your code to LLM providers through Kerno's proxy.

Kerno currently uses Anthropic and OpenAI as LLM providers, all configured with zero-data-retention tiers. Provider terms guarantee that your code is not stored, logged, or used for training or fine-tuning.

### What Kerno collects

The Kerno agent reports metadata about your test runs to Kerno's services:

* Discovered entry points (HTTP method, path, file location)
* Validation results (outcome, and counts of scenarios run, added, updated, removed, and differences detected)
* Repository and branch name

No source code is sent.

### Run reports

Kerno uploads the report of every test run your agent starts so it can be viewed in the [Kerno portal](/docs/portal/overview) by you and your team.

A run report includes the HTTP exchange each scenario produced against your application: request body and headers, response body and headers, and status code. This is what the portal's report view renders, and it is attributed to the account that ran the test.

{% hint style="info" %}
These are the request and response payloads your application produced during a test run, not your source code. If your entry points return sensitive data, that data is part of the report. Consider this when testing against environments holding production or production-like data.
{% endhint %}

### Encryption and compliance

All data Kerno manages is encrypted at rest with AES-256 and in transit with TLS 1.2+. SOC 2 Type II compliance is in progress.

For questions, contact us at <security@kerno.io>.


# Overview

Your team's shared view of what Kerno tests across your repositories, and where you set up Kerno and manage your organization.

The Kerno portal is where your whole team sees what Kerno tests across every repository. It shows which user flows and entry points have tests, every run and what it found, and how behaviour changed over time. It is also where you set up Kerno for the first time and manage your organization, members and billing.

Sign in at [portal.kerno.io](https://portal.kerno.io)

### Getting started in the portal

The first time you sign in, the portal walks you through setup:

1. **Create your organization**, and invite teammates if you want to. You can skip the invites and do it later.
2. **Learn what Kerno does**, in a short four-card introduction.
3. **Set up Kerno.** The setup page gives you a prompt to paste into your coding agent, then follows your agent's progress through five steps: Kerno connected, Codebase indexed, Environment connected, Calibration saved and User flows detected. When a step needs you, the page says so, and gives you a prompt to unblock it.
4. **Pick the user flows to test.** Choose up to 10 of the flows your agent mapped. The portal gives you a prompt that asks your agent to write and run tests for them.

Once your first baseline is recorded, the portal takes you to Analytics. Until then, a **Get Kerno running** card in the sidebar shows your setup progress and takes you back to where you left off.

### Screens

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><i class="fa-chart-line" style="color:$info;">:chart-line:</i> <strong>Analytics</strong></td><td>Coverage, code changes tested and potential issues caught across your repositories.</td><td><a href="/docs/portal/analytics">Analytics</a></td></tr><tr><td><i class="fa-list-check" style="color:$info;">:list-check:</i> <strong>Coverage</strong></td><td>Every user flow and entry point Kerno tests, and which ones have tests today.</td><td><a href="/docs/portal/coverage">Coverage</a></td></tr><tr><td><i class="fa-flask" style="color:$info;">:flask:</i> <strong>Runs</strong></td><td>Every run Kerno executed on your code, and what it found.</td><td><a href="/docs/portal/runs">Runs</a></td></tr><tr><td><i class="fa-code-compare" style="color:$info;">:code-compare:</i> <strong>Behaviour Drift</strong></td><td>How the tests behind each flow and entry point changed, period by period.</td><td><a href="/docs/portal/behaviour-drift">Behaviour Drift</a></td></tr><tr><td><i class="fa-folder-tree" style="color:$info;">:folder-tree:</i> <strong>Repos</strong></td><td>Every repository Kerno has seen, with its coverage, runs and drift.</td><td><a href="/docs/portal/repos">Repos</a></td></tr></tbody></table>

Coverage, Runs and Behaviour Drift open details in a side panel, so you can step from a run to its entry point or flow and back without leaving the page.

### Settings

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><i class="fa-users" style="color:$info;">:users:</i> <strong>Orgs and Teams</strong></td><td>Manage your organization, invite teammates, and set roles.</td><td><a href="/docs/portal/organizations-and-teams">Organizations and Teams</a></td></tr><tr><td><i class="fa-plug" style="color:$info;">:plug:</i> <strong>Integrations and API keys</strong></td><td>Issue an API key, pick up the CI workflow, and connect GitHub.</td><td><a href="/docs/portal/integrations">Integrations and API keys</a></td></tr><tr><td><i class="fa-credit-card" style="color:$info;">:credit-card:</i> <strong>Billing</strong></td><td>Manage your plan, payment method and invoices.</td><td><a href="/docs/portal/billing">Billing</a></td></tr></tbody></table>


# Analytics

The portal's home screen, a scan of coverage and testing activity across your repositories.

Analytics is the portal's home screen. It answers how testing is going across your repositories for the period you pick.

Choose a repository, or all of them, and a period of 7, 30 or 90 days. The repository you choose carries over to the other screens.

### Headline figures

A strip of tiles sums up the period:

* **Flows covered** and **Entry points covered**, the share of your user flows and entry points that have tests.
* **Code changes tested**, the code changes Kerno reviewed.
* **Potential issues caught**, the runs where a test stopped matching its baseline.
* **Tests in your suite** and **Active developers**.

### Cards

Below the tiles, a card for each of these figures shows how it moved over the period and which repositories lead or lag, plus a card of your top users by runs. Each card links to the screen with the detail behind it, such as [Coverage](/docs/portal/coverage) for flows and entry points, [Runs](/docs/portal/runs) for code changes and potential issues, and [Repos](/docs/portal/repos) for your suite.

Until Kerno has discovered your entry points, the screen points you to finish setting up instead.


# Coverage

Every user flow and entry point Kerno tests, which ones have tests today, and how to mark the critical ones.

Coverage lists the [user flows](/docs/core-concepts/user-flows) and entry points Kerno tests, and which ones have tests today, across all branches. Pick a repository to narrow it down.

### User flows

The user flows table lists each flow your agent mapped, with its number of steps and its last run. A flow counts as covered once its own flow test has finished a run, on any branch. Filter by covered or not covered, or search by name.

Select a flow to open its panel, which shows the outcome the flow should reach, its steps in order, marked as under test or setup, how its test changed over time and its run history. Select a step to open that entry point.

### Entry points

The entry points table lists every HTTP route, MCP tool and background consumer Kerno discovered. Each row shows whether the entry point has a baseline, how many user flows use it, and its last run. Filter by kind, by whether it is part of a flow, or by whether it is critical, and add columns such as the testing mode, the number of tests or the source file.

Select an entry point to open its panel, which shows its latest run, the flows it is part of, how its tests changed over time and its full run history.

### Critical entry points

Use the **Critical** switch on an entry point to mark it as critical for your team. Kerno always tests a critical entry point at high effort, whatever effort your agent asks for, and the mark is saved in your repository. An entry point Kerno no longer discovers can't be marked.


# Runs

Every run Kerno executed on your code, what each one found, and how to read a run report.

Runs lists every run Kerno executed on your code, and what it found. Pick a repository and a period of 7, 30 or 90 days.

### The run ledger

Tiles at the top count the period's runs, the ones that created a baseline, the ones that passed, and the ones with diffs. Below them:

* **Your runs in progress** shows the runs you started that are still going, with their current stage. A run waiting for you says so, and you answer it in your coding agent.
* **The ledger** lists finished runs. Switch between all runs, baselines created, passed and with diffs, search, and filter by kind (flow, HTTP, MCP or consumer), run type (validation or baseline), branch, author or test type.

### Run reports

Select a run to open its report. The report shows:

* The run's result, branch, commit, author and duration.
* The tests, grouped into diffs, updated, new, removed, unchanged and blocked. Select a test to see its details.
* For a user flow, the steps the run walked.
* For a run that stopped early, what stopped it.

When a test carries a suspected bug, the report explains it, and **Investigate with your agent** copies a prompt you can hand to your coding agent.

Your agent gives you the link to a run's report when it starts the run, so you can follow it here as it happens.


# Behaviour Drift

How the tests behind each user flow and entry point changed over time.

Behaviour Drift shows how the tests behind each user flow and entry point changed, period by period. Use it to see where behaviour moved, when, and on which branch.

Pick a repository and a period of 24 hours, 7 days or 30 days, and choose whether to show all changes, or only tests added, updated or removed.

The screen has two grids, one for user flows and one for entry points. Each row is a flow or an entry point, and each square is a slice of the period, shaded by how many tests changed in it. Hover a square for the breakdown, and select it to open the change.

The change panel lists the baselines created and the tests added, updated and removed. Expand a change to compare its before and after, and open the run report behind it.

A row appears once your agent writes, updates or removes a test for that flow or entry point.


# Repos

Every repository Kerno has seen, with its coverage, runs and behaviour drift.

Repos lists every repository Kerno has discovered entry points in or run against. Each row shows how many of its user flows and entry points are covered, its runs in the period you pick, and when it last ran. Repositories that have not run yet are grouped at the bottom.

Open a repository for an overview of its coverage, runs and behaviour drift. Pick a branch, or the latest one, and a period. The repository page brings together its headline figures, its [Coverage](/docs/portal/coverage) tables and its [Behaviour Drift](/docs/portal/behaviour-drift) grids. Opening a repository also selects it on the other screens.


# Organizations and Teams

Set up your organization, invite your team, and manage roles and access.

Your organization is your team's workspace in Kerno. It holds your members and the coverage, issues, and usage from everyone's repositories. You can belong to more than one organization and switch between them.

### One organization or several

If you belong to more than one organization, switch between them from the organization dropdown in the sidebar. You can create a new organization from the same menu. Each one keeps its own members, data, and billing, and nothing is shared between them.

### Roles

Every member has one of three roles:

* **Owner.** Full control of the organization. Only the owner can rename it or transfer ownership, and each organization has one owner.
* **Admin.** Manages members and billing, and can delete the organization.
* **Member.** Standard access to the organization's coverage, issues, and usage.

### Inviting teammates

Admins invite teammates from the **Team** page:

1. Open **Team** and select **Invite User**.
2. Enter one or more email addresses and assign each a role.
3. Send. Each invitee gets an email link that adds them to the organization once they open it and sign in.

Admins see the invites they sent under **View Invites**, where each one shows as pending or expired and can be resent or revoked.

### Managing members

The **Team** page lists everyone in the organization. Search by name or email, or filter by role, or show pending invites. From here, admins can change a member's role, update their details, or remove them from the organization. An owner can transfer ownership to another member.

### For invited members

When you accept an invite, you land in the organization and see the same coverage, issues, and usage as the rest of the team. Sign in with the account tied to your invited email, or create one if you are new to Kerno.

### Organization settings

Open the organization settings page to rename the organization and change its URL slug, or delete the organization (admins and owners). Deleting an organization removes its data for everyone, so treat it with care.

### What's next

* [Billing](/docs/portal/billing) to manage your plan and seats.
* [Analytics](/docs/portal/analytics) to see how testing is going across your repositories.


# Integrations and API keys

Create an API key for non-interactive sign-in, connect GitHub, and pick up the CI workflow from the portal.

The portal's **Settings** section is where you connect Kerno to the rest of your toolchain.

### API key

**Settings → API Key** issues your personal key for the organization you have selected, so you get one key for each organization you belong to. Choose when it expires as you create it, from one hour to never. The expiry can't be changed later, so delete the key and create a new one to change it. The card tells you when a key has expired or is inactive.

The key exists for signing in where a browser cannot open, such as CI or a remote shell. The card gives you a ready-to-copy `kerno login --api-key` command with your key in it. You can also set the key in the environment, and the CLI authenticates on its own:

```bash
export KERNO_API_KEY=<your key>
export KERNO_ORG_ID=<organization id>   # optional, when you belong to several
kerno init -w /absolute/path/to/your/repo
```

`KERNO_ORG_ID` picks the organization when your key belongs to more than one.

{% hint style="warning" %}
An API key authenticates as you. Keep it in your CI provider's secret store, and delete it from the portal if it leaks.
{% endhint %}

### Integrations

**Settings → Integrations** is where Kerno meets the places your code lives and runs.

**CI setup.** The page carries a ready-made workflow you can copy straight into your repository. It runs your existing Kerno tests on every pull request and needs no account, no API key and no agent. The full walkthrough is in [How to Run your Baseline Tests in CI](/docs/guides/run-tests-in-ci).

**GitHub.** Where it is available, connecting the Kerno GitHub App to your GitHub organization lets Kerno see the repositories you choose. Only the organization owner can connect or disconnect it. The card shows the connected account and the repositories Kerno can see. Disconnecting revokes Kerno's access, and deleting a Kerno organization revokes it as well.

### What's next

* [Analytics](/docs/portal/analytics) to see how testing is going across your repositories.
* [Organizations and Teams](/docs/portal/organizations-and-teams) to manage members and roles.


# Billing

Upgrade your plan, and manage your payment method and invoices in the Kerno portal.

Billing lives under **Settings → Billing** and is open to admins and owners. Other members don't see it in the sidebar.

### Plans

The Billing page shows the plans you can choose from: Community, which is free, Pro, and Enterprise, which is arranged through Contact Sales. For what each plan includes and current pricing, see [kerno.io/pricing](https://www.kerno.io/pricing).

### Upgrading

Select **Upgrade** on the Pro plan to start a secure Stripe checkout. Once you are subscribed, the page shows your plan as current and a summary of your next bill.

### Payment method and invoices

Under **Payment Method**, select **Edit** to open the Stripe customer portal, where you manage your payment method. Once you have invoices, they are listed on the Billing page, and each one opens its receipt.

### What's next

* [Organizations and Teams](/docs/portal/organizations-and-teams) to add the teammates your subscription covers.
* [Analytics](/docs/portal/analytics) to see how testing is going across your repositories.


# Installation Issues

Solutions to Kerno installation issues.

### The agent is not starting

* **Turn off your VPN.** If a VPN is intercepting outbound traffic, Kerno cannot reach LLM providers and will fail to index your codebase or run tests.
* **Check the startup log.** `~/.kerno/agent.stdout.log` records what happened when the agent tried to boot, and is the fastest way to see the real error.
* **Sweep orphaned processes.** Run `kerno doctor --clean` to kill stray agent processes and clear inconsistent state files.
* **Start clean.** `kerno reset` stops applications, clears workspace caches, and restarts the agent fresh.

### My coding tool can't reach Kerno after a restart

Kerno keeps the same MCP port across restarts when it can, so your coding tool keeps working after `kerno reset`, `kerno stop` or `kerno init --force-switch`. If another program took the port, `kerno init` prints a new URL. Update your coding tool with it and reconnect.

### `kerno init` says the agent is bound to another workspace

The agent serves the directory `kerno init` started it for, and every directory nested under it. `kerno init` matches the exact path, so running it from another repository, or from a subdirectory of the bound one, reports this conflict. Run `kerno init` from the bound directory to reuse the agent.

To switch, answer yes at the prompt in a terminal. In a non-interactive shell, such as inside a coding agent, run `kerno stop` first or pass `--force-switch`:

```bash
kerno init -w /absolute/path/to/other/repo --force-switch
```

### The bare `kerno` shell does nothing inside my coding agent

Running `kerno` with no command opens an interactive terminal UI, which needs a real terminal.

From inside a coding agent, run `kerno status`. It checks sign-in, Docker, the agent and your apps, then exits, with an error when a check fails. Ask your agent to run a Kerno healthcheck for everything else.

### The port isn't in `~/.kerno/mcp.port`

Those files exist only while the agent is running. If `agent.pid`, `agent.port` or `mcp.port` are missing, the agent is not up. Check `~/.kerno/agent.stdout.log`.

### The indexing process is failing

* Make sure your [language stack is supported](/docs/references/supported-technologies#supported-backend-languages).
* Kerno indexes git-tracked files. Brand-new files that have never been added to git are not analyzed until you stage or commit them.
* Try `kerno reset` to clear cached analysis and start fresh.

### Collecting logs for support

```bash
kerno export-logs -o ./kerno-logs.zip
```

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we'll gladly help.
{% endhint %}


# Environment Setup Issues

Solutions to Kerno test environment setup issues.

### Kerno can't reach my application

Kerno does not start your application. You do. When setup fails, it is almost always because the URL you gave Kerno is not reachable from inside Kerno's sandbox.

* **Confirm your app is actually running** and responds on the URL you configured. Try it with `curl` first.
* **Give Kerno the URL you use on your machine**, such as `http://localhost:3000`. Kerno's sandbox routes `localhost` to your machine by itself.
* **Check the readiness signal, not the container state.** Ask your agent for the environment status and look for `ready_for_endpoint_test`. Containers being up is not the same thing.
* If your app's URL or credentials changed, save the configuration again.
* **Force a network mode if your app is still unreachable.** Kerno picks how its sandbox reaches your machine based on your Docker engine. To override it, stop Kerno and start it again with `KERNO_SANDBOX_NETWORK` set:

  ```bash
  kerno stop
  KERNO_SANDBOX_NETWORK=host kerno init
  ```

  `host` puts the sandbox on your machine's network and gives up its network isolation. `gateway` routes through `host.docker.internal`, the default on macOS, Windows and Docker Desktop. `dnat` forces the native Linux routing.

### Kerno says it can't derive my database schema

Direct database access needs two things: credentials, and a schema Kerno can derive from your source code. Credentials alone are not enough.

Kerno finds schemas by filename and path convention, covering Prisma schemas, `schema.rb`, SQL files under a migrations directory, Drizzle, Alembic, Liquibase, and numbered migration files. Projects using Knex, Sequelize, or TypeORM with descriptively-named migration files are often not recognised.

When the plan has tests that need direct database access, Kerno pauses and asks you. Answer with a repository-relative path to your schema file, or paste the DDL directly. Kerno saves it for that entry point and reuses it on later runs.

If you reply `skip`, Kerno blocks the tests that need the database and does not ask again for that entry point. Its other tests still run.

### A scenario reports as blocked

Blocked means the scenario never ran, usually because a dependency it needs is not configured. It is neither a pass nor a fail, and nothing was tested. Configure the missing dependency and run again.

### Running Docker without root

Docker must be usable without `sudo`. This is required for Kerno's agents that interact with Docker automatically.

Add your user to the [Docker group](https://docs.docker.com/engine/install/linux-postinstall/#manage-docker-as-a-non-root-user), you will need to launch a new terminal to make it works:

```bash
sudo usermod -aG docker $USER
# Testing
docker run hello-world
```

### EOF failure during analysis

**Symptom:** Kerno fails during the initial analysis phase with an error containing `EOF failure` (e.g. `AnthropicLLMClient — EOF failure`).

**Cause:** A VPN is blocking Kerno’s outbound connection to the LLM API.

**Fix:** Disable your VPN and restart the Kerno agent.

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we’ll gladly help.
{% endhint %}


# Test Execution Issues

Solutions to test execution issues.

### My test run looks stuck

The most common cause is not a stuck run at all: it is a run waiting for you.

When Kerno pauses for plan approval or an answer, the run reports that it needs your input and makes no progress until you reply. Ask your agent for the pending question, and answer it.

If nothing is pending, ask your agent to read the run's state directly with `kerno_get_state`, or to cancel and start again.

### Kerno says there's nothing to validate

Validation runs the scenarios already on disk for an entry point. If none exist, the run stops immediately and tells you to generate first. Nothing is created implicitly.

Generate tests for that entry point before validating it.

If you asked "what did my changes affect?" and got nothing back, check whether your changes are still uncommitted. The `changed` scope compares your working tree against `HEAD`, so once everything is committed there is nothing in the diff. Target the entry point or file directly instead.

### Kerno says my entry point doesn't exist

When you name a route Kerno's analysis did not list, Kerno searches your code for its handler before refusing.

* If the handler serves a different path, Kerno names that path. Ask again with the path your app serves.
* If no handler is found, Kerno lists the closest routes it knows. Check the method and path against your app, then ask again.
* If Kerno found the route but not its HTTP method, name the method yourself:

```
Use Kerno to generate tests for POST /users.
```

### Tests are failing with 401s or 403s

Kerno works out how each entry point authenticates by reading your source. When it gets this wrong, every scenario fails at the door.

* Check whether your auth needs a credential Kerno cannot obtain by itself. If your app signs tokens with a secret from an environment variable, provide it so Kerno can construct a valid token.
* Tokens Kerno signs itself use HMAC algorithms. If your verifier requires `RS256` or another asymmetric algorithm, Kerno needs a real login route to call instead.
* Flows requiring a person to complete a hand-off with an external identity provider cannot be automated. Those scenarios report as blocked.

### A scenario reports as blocked

Blocked means the scenario could not run, usually because a dependency it needs is not configured. It is neither a pass nor a fail. See [Environment Setup Issues](/docs/troubleshooting/environment-setup-issues).

### A scenario reports as not implemented

Kerno could not produce a working scenario for that case. This is reported honestly rather than counted as a pass. Try running an update on the entry point with guidance describing what that scenario should do.

### Kerno flagged a "potential bug" that isn't one

Potential bugs attach to scenarios that pass: Kerno found the entry point doing something the plan did not expect and documented the real behaviour.

If it is known or intended, tell your agent to ignore it and give your reason. Kerno records the decision for that entry point, and later reports, validation runs included, stop showing it. The test file keeps its potential-bug note until you next update that entry point's tests, which then document the behaviour as expected. Note that there is currently no way to reverse this through Kerno.

### My entry point changed on purpose and now everything fails

That is working as intended. Run an update rather than regenerating:

> "Use Kerno to update the tests for POST /users. The response no longer includes the role field."

Update keeps every scenario that still applies and edits only what your change requires.

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we'll gladly help.
{% endhint %}


# Kerno on Arch Linux

Configure Docker on Arch Linux so Kerno can run.

This page explains how to configure your development environment on Arch Linux to run Kerno. Kerno relies on Docker to run its test sandbox and to reach your application.

### Docker Desktop on ArchLinux

Docker Desktop is not officially supported on ArchLinux. It may fail with errors such as:

```
qemu: process terminated unexpectedly: signal: aborted (core dumped)
```

These issues are commonly related to kernel updates, QEMU integration, or unsupported system configurations. For ArchLinux, Docker Engine installed directly from the distribution repositories is the recommended and supported approach.

{% hint style="info" %}
We highly recommend uninstalling Docker Desktop and install the deamon directly instead.
{% endhint %}

### Docker installation on ArchLinux

Docker must be installed using the official Arch Linux packages. Following the Arch [Wiki documentation](https://wiki.archlinux.org/title/Docker):

```bash
sudo pacman -Syu
sudo pacman -S docker
sudo systemctl enable --now docker
```

You can verify the installation with:

```bash
docker info
```

### Docker Compose

Kerno requires Docker Compose v2, which uses the following syntax:

```bash
docker compose
```

The legacy `docker-compose` command is deprecated and should not be used. Installation instructions for the `docker compose` plugin are available in the official [Docker documentation](https://docs.docker.com/compose/install/linux/), follow the manual installation, since ArchLinux distro isn't supported.

After installation, verify with:

```bash
docker compose version
```

### Verification checklist

Before using Kerno, ensure the following commands work without `sudo`:

```bash
docker info
docker run hello-world
docker compose version
```

If these commands succeed, your system is correctly configured to run Kerno end-to-end tests.

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we’ll gladly help.
{% endhint %}


# Test Kerno with OSS Projects

Open source projects you can use to quickly try Kerno and see how it behaves in a real codebase.

You can use one the open source projects listed below to try Kerno in a safe and simple way. This will give you a clear understanding of how Kerno works and how it integrates into your own workflow.

### How You Use a Sample Project

1. Fork a project from the list.
2. Clone it to your machine.
3. Start the backend locally, and note the URL it listens on.
4. Open the project in your editor with [Kerno installed](/docs/getting-started/quickstart).
5. Point Kerno at the running app, and give it the Postgres connection string.
6. Generate tests for an entry point, then change the code and validate to see how Kerno reacts.

This lets you understand the full workflow without any risk.

### Project List

{% tabs %}
{% tab title="Typescript/JavaScript" %}

<table><thead><tr><th>Project Name</th><th data-type="content-ref">GitHub Repo</th><th>Framework</th><th>Dependencies</th></tr></thead><tbody><tr><td>Typescript Example Project</td><td><a href="https://github.com/kernoio/example-typescript">https://github.com/kernoio/example-typescript</a></td><td>Express.js</td><td>PostgreSQL</td></tr></tbody></table>
{% endtab %}

{% tab title="Python" %}

<table><thead><tr><th>Project Name</th><th data-type="content-ref">GitHub Repo</th><th>Framework</th><th>Dependencies</th></tr></thead><tbody><tr><td>Python Example Project</td><td><a href="https://github.com/kernoio/example-python">https://github.com/kernoio/example-python</a></td><td>FastAPI</td><td>PostgreSQL</td></tr></tbody></table>
{% endtab %}
{% endtabs %}

{% hint style="info" %}
If you encounter issues or have questions, [message us on Slack](https://join.slack.com/t/kerno-community/shared_invite/zt-3fzrjxgog-voatSyyKY78uDj6QaNW4OQ), and we’ll gladly help.
{% endhint %}


