Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 22 additions & 5 deletions .env.template
Original file line number Diff line number Diff line change
Expand Up @@ -10,13 +10,11 @@
KONTENT_CLI_AMPLITUDE_API_KEY=

# Base Kontent.ai domain; all endpoint URLs derive from it (app.<domain>, manage.<domain>).
# Baked into the build from here; defaults to kontent.ai (production). Use
# devkontentmasters.com for dev. A runtime KONTENT_URL in your shell overrides it.
# Baked into the build from here; defaults to kontent.ai (production).
KONTENT_URL=devkontentmasters.com

# Auth0 tenant baked into the build; defaults here target the QA/dev tenant.
# A runtime KONTENT_AUTH0_* var in your shell overrides each. The requested
# scope is not configurable (fixed requirement of the CLI).
# A runtime KONTENT_AUTH0_* var in your shell overrides each.
KONTENT_AUTH0_DOMAIN=login.devkontentmasters.com
KONTENT_AUTH0_CLIENT_ID=
KONTENT_AUTH0_AUDIENCE=https://app.kenticocloud.com/
Expand All @@ -25,7 +23,6 @@ KONTENT_AUTH0_AUDIENCE=https://app.kenticocloud.com/

# Telemetry opt-out (set to 1/true to disable). DO_NOT_TRACK is the standard equivalent.
KONTENT_DO_NOT_TRACK=
DO_NOT_TRACK=

# Verbose telemetry logging for debugging.
KONTENT_TELEMETRY_DEBUG=
Expand All @@ -43,3 +40,23 @@ E2E_SOURCE_ENV_ID=
# Domain of the e2e project. The e2e run ignores KONTENT_URL above and defaults
# to production kontent.ai; set this only when the test project lives elsewhere.
E2E_KONTENT_URL=kontent.ai

# --- Agent evals (pnpm evals:run; fails with an error when these are unset) ---
# Read from .env by vitest.evals.config.ts, same as the E2E_* block above.

# Management API key of the dedicated evals project: access to all environments
# plus the Manage environments permission.
EVALS_MAPI_KEY=

# Empty template environment that each eval run clones. Never written to.
EVALS_SOURCE_ENV_ID=

# Model under test, passed to the Agent SDK. Defaults to "sonnet" when unset.
EVALS_MODEL=

# Set to 1 to keep the cloned environment after the run instead of deleting it.
EVALS_KEEP_ENV=

# Evals authenticate as your local Claude Code subscription login; the run
# fails fast if ANTHROPIC_API_KEY is set. Set this to 1 to run with a key anyway.
EVALS_ALLOW_API_KEY=
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -13,5 +13,9 @@ my-kickstart/
.claude/worktrees/
.claude/plans/

# Eval run reports; the folder itself stays tracked
evals/results/*
!evals/results/.gitkeep

# pnpm pack output
*.tgz
13 changes: 8 additions & 5 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,15 +23,16 @@ Three layers, dependencies point downward only (`commands → core → lib`):
- `src/core/**` — orchestration of business logic. Returns `Result`/`Option`; never writes to the console directly (logs only through a passed `Logger`). **Exception:** interactive commands may drive their own terminal UI from core — e.g. `src/core/project/bootstrap.ts` uses the prompts of `src/lib/ui/prompts.ts` (spinners, `confirm`/`select`, notes) directly because the flow is inherently interactive. Keep non-interactive core free of direct console writes.
- `src/lib/**` — reusable primitives: `auth/`, `iapi/`, `mapi/`, `config/`, `telemetry/`, plus `result.ts` and `option.ts`.

Adding a command: export a `register: RegisterCommand` (see `src/commands/login/login.ts`), then add its import to the `register` array in the parent command or `src/commands/registry.ts`. Then run `pnpm docs:generate` (`scripts/generateCommandDocs.ts`) — it replays the registrations against a recording proxy and rewrites the generated docs: the marker-fenced command table in the root `README.md`, and the `<!-- reference:start/end -->` block in each command folder's `README.md` (created as a skeleton when missing). Prose outside the markers is handwritten — write command docs there, never inside the block. Two opt-out sets in the script: `commandsWithoutPage` (no colocated README) and `commandsWithoutIndexEntry` (no root-README table row; telemetry is there). The generator errors on a command-folder README with markers but no matching command (stale after rename/removal) — resolve by hand; it never deletes pages.
Adding a command: export a `register: RegisterCommand` (see `src/commands/login/login.ts`), then add its import to the `register` array in the parent command or `src/commands/registry.ts`. Then run `pnpm docs:generate` (`scripts/generateCommandDocs.ts`) — it replays the registrations against a recording proxy and rewrites the generated docs: the marker-fenced command table in the root `README.md`, and the `<!-- reference:start/end -->` block in each command's `README.md` (created as a skeleton when missing), placed in the leaf's own folder when it has one and in its parent group's folder otherwise. Prose outside the markers is handwritten — write command docs there, never inside the block. Two opt-out sets in the script: `commandsWithoutPage` (no colocated README) and `commandsWithoutIndexEntry` (no root-README table row; telemetry is there). The generator errors on a command-folder README with markers but no matching command (stale after rename/removal) — resolve by hand; it never deletes pages.

### API clients

- `iapi` (`src/lib/iapi`) — internal Kontent.ai API; hand-rolled client, one file per endpoint, over `@kontent-ai/core-sdk`. Endpoint validators (the `schema` field) must be **`zod/mini`** (`import * as z from "zod/mini"`) — classic `zod` won't infer the payload.
- `mapi` (`src/lib/mapi`) — public Management API via `@kontent-ai/management-sdk`. `src/lib/mapi/raw` is the deliberate opposite: a passthrough (no schema, no response interpretation) behind `kontent mapi`, where a 4xx/5xx is a result, not an error. It builds on core-sdk's `getDefaultHttpService` and turns the non-2xx it reports as errors back into results, reading the body off `error.details.adapterResponse`; retry, `Retry-After` and header merging are core-sdk's. Its doc comments carry the why: `raw/client.ts` for which SDK error reasons stay errors, `raw/contentType.ts` for the rule that decides whether a body is printed.
- `@kontent-ai/core-sdk` — shared HTTP/SDK layer both clients build on.
- `learn` (`src/lib/learn`) — tokenless `createFetchQuery` client over the Learn-MCP service (`https://learn-mcp.kontent.ai`, no auth, GET only) behind `kontent docs`. The `zod/mini` schemas declare only the fields the CLI reads; core-sdk hands back the raw payload, so unknown keys survive to stdout. `runtimeValidation.validateResponses` must stay on, otherwise the schema never runs.
- `@kontent-ai/core-sdk` — shared HTTP/SDK layer all clients build on.

**Commands build clients; core receives them.** The command builds the `iapiClient`/`mapiClient` and passes them into core (e.g. `performBootstrap(params, { logger, iapiClient, mapiClient })`); core never constructs clients itself. Auth failure is handled in the command, not surfaced as a core `Result` error. Same split for arguments: pure parsers live in `lib` (`mapi/raw/headers.ts`, `mapi/raw/method.ts`), reading what the invocation points at stays in the command, and each layer declares only the error kinds it raises.
**Commands build clients; core receives them.** The command builds the `iapiClient`/`mapiClient` and passes them into core (e.g. `performBootstrap(params, { logger, iapiClient, mapiClient })`); core never constructs clients itself. Auth failure is handled in the command, not surfaced as a core `Result` error. The `kontent login` token is the credential for both `iapi` and `mapi`: a logged-in user needs no separate API key, so only the environment id ever needs resolving. Same split for arguments: pure parsers live in `lib` (`mapi/raw/headers.ts`, `mapi/raw/method.ts`), reading what the invocation points at stays in the command, and each layer declares only the error kinds it raises.

### Output channels

Expand All @@ -54,15 +55,17 @@ Core takes the logger as a parameter or inside its `deps` object; `createLoggerF
- **ESM import extensions:** relative imports must end in `.js` (biome enforces `useImportExtensions`).
- **No redundant wrappers.** Don't add a function that only forwards to another; reuse existing helpers instead of duplicating logic.
- **Comments only for non-obvious "why".** No restating-the-code comments, no repeating a fact already stated elsewhere (put domain facts once, on the type), no justifying a change to the reviewer ("X already did Y, so..."). No emojis anywhere.
- **Exports first.** A module's exported types and functions go at the top, private helpers below them. Arrow consts are only called after module evaluation, so referring downward is safe.
- **Exports first.** A module's exported types and functions go at the top, private helpers below them. Below them, private helpers in call order, depth-first; a private constant sits directly above its one user. Arrow consts are only called after module evaluation, so referring downward is safe.
- **No barrel files** except a deliberate public API.

## Testing

Vitest; `test/unit/` for pure unit tests, `test/integration/` for integration tests, `test/helpers/` for shared helpers. Command-level behavior (argument parsing, exit codes, which stream a message lands on) is tested by folding a command's `register` over a real yargs instance and faking only the core call underneath — see `test/integration/mapiCommand.test.ts`. Run `pnpm test`. Inject fakes into core instead of real I/O — for iapi reuse `test/helpers/iapiTestClient.ts` (real client over core-sdk's `HttpAdapter` seam, declarative routes).
Vitest; `test/unit/` for pure unit tests, `test/integration/` for integration tests, `test/helpers/` for shared helpers. Command-level behavior (argument parsing, exit codes, which stream a message lands on) is tested by folding a command's `register` over a real yargs instance and faking only the core call underneath — see `test/integration/mapiCommand.test.ts`. Run `pnpm test`. Inject fakes into core instead of real I/O — for iapi reuse `test/helpers/iapiTestClient.ts` (real client over core-sdk's `HttpAdapter` seam, declarative routes). Harness unit tests live in `evals/test/` (helpers in `evals/test/helpers/`) and run in `pnpm test`; only the paid agent run is opt-in.

`test/e2e/` runs the built binary against a real Kontent.ai project (clone-per-run from an empty template env). Gated on `E2E_MAPI_KEY`/`E2E_SOURCE_ENV_ID` (fails fast with an error when unset). Run with `pnpm test:e2e` (own `vitest.e2e.config.ts`, loads `.env`); excluded from `pnpm test` and the before-halting gate. CI: `.github/workflows/e2e.yml` (master push, PRs, manual; fork PRs are skipped at the job level — no secret access).

`evals/` is a Vitest-driven agent-eval harness (`pnpm evals:run`, own `vitest.evals.config.ts`): the Agent SDK drives the built CLI with Bash (confined to a workspace dir) plus WebFetch (`kontent.ai` only, key-in-URL denied), against one cloned environment shared by every task, run sequentially in dependency order (a task whose parent did not PASS is recorded BLOCKED, no agent spawned). The policy is a pure reducer (`evals/lib/policy.ts`) wired into the `PreToolUse` hook by `evals/lib/agent.ts`, the only effectful module; tool calls and denials are derived from the run's raw messages in `evals/lib/toolCalls.ts`; CLI invocations and exit codes come from the shim's per-task log (`evals/lib/invocations.ts`), never from parsing shell text. The agent run never runs in `pnpm test` or CI, and is gated on `EVALS_MAPI_KEY`/`EVALS_SOURCE_ENV_ID` (which the CLI itself must never read). `evals/tasks/<id>/task.ts` (prompt + check together) grades the resulting state through the Management SDK; checks stay deterministic: return `Result` (a failed lookup is an `err`, never a failed assertion), and stay tolerant about names the task did not fix. Reports (`evals/lib/report/`, markdown only, no LLM calls) are a nice-to-have removable by deleting that folder and its one call site. Playbook: `evals/README.md`.

## Telemetry

Amplitude-based, see `TELEMETRY.md`. Env vars are read from `process.env` where they apply, never mapped onto yargs options — `src/index.ts` deliberately does not call `.env()`, so a stray `KONTENT_*` var cannot break an unrelated command.
Expand Down
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,10 +56,13 @@ Each command supports `--help` for its own options.
<!-- commands:start -->
| Command | Description |
| --- | --- |
| [`kontent docs search <query>`](src/commands/docs/search/README.md) | Search Kontent.ai Learn docs and API reference, ranked by relevance |
| [`kontent docs endpoint <query>`](src/commands/docs/endpoint/README.md) | Show matching API endpoints: method, URL, parameters, responses, code samples |
| [`kontent docs object <query>`](src/commands/docs/object/README.md) | Show matching API reference objects and their properties |
| [`kontent login`](src/commands/login/README.md) | Authenticate with Kontent.ai via Auth0 device flow |
| [`kontent logout`](src/commands/logout/README.md) | Clear stored authentication tokens |
| [`kontent mapi <endpoint>`](src/commands/mapi/README.md) | Send an authenticated request to the Management API |
| [`kontent project sample bootstrap`](src/commands/project/sample/README.md) | Clone a sample app for an environment and wire its .env |
| [`kontent project sample bootstrap`](src/commands/project/sample/bootstrap/README.md) | Clone a sample app for an environment and wire its .env |
<!-- commands:end -->

## Global options
Expand Down
83 changes: 83 additions & 0 deletions evals/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# Agent evals

Measures how far an AI agent, acting like a fresh human user, gets when it drives the built
`kontent` CLI against a real Kontent.ai environment. One Agent SDK run per task; deterministic
checks in `evals/tasks/<id>/task.ts` grade the resulting environment state afterwards. This agent
run goes through Vitest, but not as part of `pnpm test` or CI - it costs money and needs a cloned
environment, so it is a deliberate opt-in. The harness's own unit tests (`evals/test/`) are plain
Vitest specs and run as part of `pnpm test` like any other unit test.

## What the agent gets

Two tools, declared in `evals/lib/policy.ts` and wired into the SDK options by `evals/lib/agent.ts`,
enforced by the `PreToolUse` hook via the pure policy reducer in `evals/lib/policy.ts`:

- `Bash`, denied for any command referencing a path outside the task's workspace directory. This is
a file-access guard, not a sandbox: Bash has full network access, and the agent is trusted with it.
- `WebFetch`, allowed for `kontent.ai` only (the API reference lives at `kontent.ai/learn/...`), and
denied outright for a URL that contains the Management API key. The rule exists to measure, not to
contain: every fetch is a fallback the CLI's own docs did not cover.

The preamble says nothing about the web or about `kontent docs`: which route the agent takes to the
API reference is part of what a run records (`docs` and `fetch` counts, and the report's "Web
fetches" section). Every fetch is a place where the CLI's own docs did not carry the agent.

## Preconditions

- `EVALS_MAPI_KEY` and `EVALS_SOURCE_ENV_ID` exported in your shell, or set in `.env`.
- Logged into Claude Code locally (`claude` on PATH, authenticated).
- `ANTHROPIC_API_KEY` unset. Evals authenticate as your local subscription login, not an API key -
see [Anthropic policy](#anthropic-policy) below. The run fails fast if it is set, unless you pass
`EVALS_ALLOW_API_KEY=1`.

## Run it

```
EVALS_MODEL=sonnet pnpm evals:run
```

`globalSetup.ts` builds the CLI, clones `EVALS_SOURCE_ENV_ID`, and hands the environment id and the
built CLI's bin directory to the test file. `evals/run.eval.ts` then runs every task exported from
`evals/lib/registry.ts`, in dependency order, sequentially, against that one environment: if any of
a task's parents did not PASS, the task is recorded BLOCKED and no agent is spawned for it. Teardown
deletes the cloned environment unless you keep it (see below).

### Knobs (environment variables)

- `EVALS_MODEL` - model passed to the Agent SDK. Default `sonnet`.
- `EVALS_KEEP_ENV=1` - keep the cloned environment instead of deleting it at the end.
- `EVALS_ALLOW_API_KEY=1` - run even with `ANTHROPIC_API_KEY` set.

## Output

Each run writes to `evals/results/<YYYY-MM-DD>-<HHMM>-<model>/` (git-ignored; date and time are UTC,
taken at the start of the run):

- `run.json` - header (model, tools, permission rules, max turns, CLI version, git sha, environment
id, preamble hash, timing) and one row per task (verdict, turns, tool call, failed, denied, help,
docs and fetch counts, cost, duration, stop reason).
- `report.md` - the same summary as a table, plus a friction section (every failed call, every
web fetch) and every task's final agent reply. Nice-to-have:
`evals/lib/report/` renders it from the same data as `run.json`, with no LLM calls; deleting that
folder and the one call into it from `evals/lib/results.ts` removes the feature cleanly.
- `tasks/<id>.json` - the task's full trace (tool calls and fetches, agent text, numbers) plus its
verdict and assertions.
- `tasks/<id>.raw.jsonl` - every raw Agent SDK message for that task, one per line.
- `tasks/<id>.md` - the same task report rendered as markdown.
- `tasks/<id>.invocations` - one JSON line per `kontent` process the agent ran (exit code and
argv, key redacted), the source of the `cli`, `failed`, `help` and `docs` columns.

Each task also gets its own workspace directory under the OS temp dir (`kontent-eval-<task>-*`),
where the agent runs its commands. These are left in place after the run, not cleaned up, so they
stay available for inspection.

## Adding a task

1. `evals/tasks/<id>/task.ts` - export an `EvalTask` with an `id`, its `dependsOn` (the ids of its
parent tasks), a `prompt` phrased as a busy human would type it, with no hints about the CLI, and
a `check` that reads the environment through the Management SDK and returns assertions. A failed
lookup is an `err`, never a failed assertion. Codenames the prompt fixes are asserted exactly;
names are never asserted; content items are named only and found by name. Include a regression
assertion for anything an earlier task in the chain seeded.
2. Add its id to the `TaskId` union in `evals/lib/types.ts`.
3. Add it to `evals/lib/registry.ts`.
Loading
Loading