Architecture
Cerbie runs entirely on Cloudflare. The goal is the best review in the shortest wall-clock time. Every design choice below either cuts latency (parallelism, warm mirrors, no installs, prompt caching) or raises signal (specialists, independent verification, convergence across rounds).
flowchart TD GH[GitHub webhooks] --> W[Worker: verify + gate] W --> DO[PullRequestCoordinator<br/>one Durable Object per PR] DO -->|debounce, supersede| WF[ReviewWorkflow] WF --> CTX[Context step] CTX --> S1[Summary] CTX --> S2[Primary reviewer<br/>sharded for large PRs] CTX --> S3[Specialists:<br/>security · data · contracts] CTX --> S4[Custom checks] CTX --> S5[Deterministic:<br/>deps · lockfile · secrets] CTX --> S6[Follow-up of<br/>earlier findings] S2 & S3 & S4 & S5 & S6 --> V[Verifier<br/>fresh context] V -->|risky PR that would be approved| C[Approval challenge<br/>+ verification of its findings] V --> P[Publish: review, inline comments,<br/>check runs, thread housekeeping] C --> P S1 -.->|posted as soon as ready| GH2[Sticky summary comment] S2 & S3 & V -.->|tools| GIT[(Git mirror container<br/>shallow bare clones)] S2 & S3 & V -.->|LLM calls| LLM[cerbie-llm-gateway<br/>CLIProxyAPI container] P --> D1[(D1: runs, findings,<br/>users, settings)] WF --> R2[(R2: context, results,<br/>rendered output)]
Request flow
Section titled “Request flow”- Webhook (
apps/worker/src/index.ts,webhooks.ts). The Worker verifies the HMAC signature and deduplicates by delivery ID in D1. It then reads.cerbie/config.yamlfrom the base branch to decide whether this event should trigger a review (drafts, ignored authors and labels, base branches,cerbie: skip). Commands (@cerbieai …) and thread replies are handled here too. - Coordinator (
coordinator.ts). One Durable Object per PR debounces bursts of pushes using a single alarm. If a newer head arrives while a review is running, the old Workflow is terminated (no tokens spent on stale code) and the new head is reviewed. Re-delivered or duplicate events for an already-reviewed head are ignored. - Workflow (
review/workflow.ts). A Cloudflare Workflow makes every stage durable and independently retried:start: creates theCerbie (Review)check run (in progress) and, on a PR’s first review, a placeholder summary comment.context: fetches PR metadata, files and patches, CI status, discussion, the linked Linear ticket, guideline files,.cerbie/, and the PR’s earlier findings. In parallel it syncs the git mirror to the exact base and head commits. The bundle goes to R2. Triage (deterministic, zero latency) picks the tier and streams.- Fan-out, all in parallel: summary, reviewer streams, custom checks, deterministic checks, follow-up of earlier findings. The summary is posted the moment it is ready.
verify: candidates are merged and deduplicated. Duplicates of still-open earlier findings are dropped so threads aren’t repeated. A verifier with a fresh context confirms, adjusts or rejects every critical and should-fix candidate against the code. The confidence floor and suggestion cap apply.challenge: only when the PR is risky (deep tier, or any specialist ran) and the verified findings would still approve it. One reviewer gets the first pass’s notes, files read and findings, and looks for failure modes nobody traced. Its should-fix and critical findings go through a second verification (verify challenge). If the challenge fails, coverage is incomplete and the PR is not approved.publish: the head guard skips publication if the PR moved. Cerbie then posts one review (verdict, inline comments with suggestions) and completes the check runs (one per custom check). It leaves earlier inline comments untouched and never resolves threads (that is the PR author’s call); for a fixed finding whose thread someone replied to, it posts a short “✅ Addressed in …” reply. It also dismisses stale Cerbie blocks and updates the summary with the verdict.
- Replies (
reply-workflow.ts). A human reply on a Cerbie thread, or@cerbieai <question>, runs a short agent with the same tools. The agent re-reviews the finding(s) a reply concerns with the reply as context: it may withdraw, confirm a fix, accept a reasoned deferral (should-fix and suggestions only) or uphold the finding, and one top-level comment can cover several findings. It is capped at three replies per thread. Its decisions go to the PR’s coordinator, which applies them (applyDecisions) and re-decides the verdict (reassessVerdict) in one serialised step, never concurrently with a review; if nothing blocks, it approves without re-reviewing the PR. Withdrawn and deferred findings are not raised again unless they come back more severe. Hand-resolved threads do not change findings.
Local reviews and the review cache
Section titled “Local reviews and the review cache”The Cerbie CLI runs the same ReviewWorkflow before a pull request exists:
cerbie loginuses GitHub device flow through the CerbieAI app (/cli/v1/login/*). The Worker keeps only a hash of the CLI session token, and checks every request against the user’s permission on the repository, matched on their immutable GitHub ID.cerbie reviewuploads a thin git bundle (merge base..head; uncommitted work becomes a temporary commit) to/cli/v1/repos/:owner/:name/reviews. The Worker reserves the user’s rate-limit quota, stores the bundle in R2 and starts the workflow withparams.local(PR number 0, never published).- The context step syncs the merge base from GitHub, imports the bundle into the git mirror and diffs there. Policy comes from the tip of the target branch on GitHub, as for a PR. Git objects are content-addressed, so the reviewed commit is exactly the code its SHA names.
Follow-up reviews are incremental (files changed since the last reviewed commit) unless decideReviewScope (packages/core/src/scope.ts) says otherwise. It uses signals that need no model call: merge base, a hash of the repository’s review rules, whole-PR triage, delta share, rounds and age since the last full review, and relative importers of changed exports, read from the git mirror. Each run’s context records the rules hash and whole-PR triage for the next decision; eval/scope-replay-2026-10-08.md replays production history through it.
Every run records a cache key: merge base, head tree and a policy hash (.cerbie/ and guideline files, org instructions, the linked ticket as read, recipe, prompts). A PR’s first review (or any local review) that finds a complete, non-incremental run with the same key from the last 14 days publishes that run’s findings instead of reviewing again; only the summary is written fresh. If an older run of the same code is still in progress, the workflow sleeps and waits for it. Runs only wait for older runs, so two can never wait on each other.
Review tiers
Section titled “Review tiers”| Tier | When | Streams |
|---|---|---|
| light | Docs, tests or lockfiles only, or ≤60 substantive changed lines with no risk signals | One focused reviewer |
| standard | Everything else | Primary reviewer, plus specialists whose signals fire |
| deep | Configured paths.risk match, >600 substantive lines, or both security and data signals |
Primary (sharded across up to four reviewers for big diffs) plus every signalled specialist |
Signals come from paths (auth, migration, routers/, …) and added code ($transaction, dangerouslySetInnerHTML, exec(, removed exports, …). Triage errs towards running a specialist: a parallel call that finds nothing costs far less than a missed lens. The deterministic checks, custom checks and follow-up always run when relevant.
Why agents don’t run code
Section titled “Why agents don’t run code”Installing dependencies for the monorepo takes minutes and gigabytes, and CI already runs type checks and tests. Cerbie reads CI status instead and spends its budget on what CI can’t catch: logic, authorisation, data integrity, concurrency and contracts. Agents get a read-only toolset over exact commits:
| Tool | Backed by |
|---|---|
read_file, list_directory |
git mirror (git show/ls-tree); falls back to the GitHub contents API |
grep, find_files |
git grep/ls-tree against the commit object, with no checkout (full monorepo grep takes about 0.5s) |
diff_file, git_log |
PR patches or git diff; shallow history |
The git mirror (containers/git) is a small Bun server. It keeps one shallow, bare clone per repository in a container that stays warm for 30 minutes. Cold sync of the monorepo takes about 5s, a warm incremental sync about 10ms. It receives a per-request, repository-scoped, read-only installation token and never executes repository code.
Models
Section titled “Models”Every LLM call goes through cerbie-llm-gateway, a CLIProxyAPI instance in a Cloudflare Container reached over a service binding. Model assignments are a recipe (apps/worker/src/llm.ts). Admins can override it in the dashboard without a deploy.
| Role | Default | Why |
|---|---|---|
| Every role defaults to GPT-6.1 Sol; always prefer it over GPT-6 Sol. Roles differ in effort and step budget. |
| Role | Default | Why |
|---|---|---|
| primary | GPT-6.1 Sol, high effort | Deepest tracing of behaviour |
| challenge | GPT-6.1 Sol, high effort | Runs after verification, only on risky PRs (deep tier or any specialist) whose verified findings would still approve. It reads the first pass’s notes and findings and hunts for failure modes nobody traced. Its findings are verified separately; if it fails, the review is incomplete rather than approved |
| specialist | GPT-6.1 Sol, high effort | Focused lenses in parallel |
| verifier | GPT-6.1 Sol, high effort | Re-checks each finding in a fresh context with its own tools |
| light | GPT-6.1 Sol, high effort | Small PRs still get the strongest reviewer |
| summary, reply | GPT-6.1 Sol, low / medium effort | Speed |
The large shared context (PR brief, guidance, numbered diff) is the first message in each conversation, so every tool-loop step reuses it from the prompt cache (automatic on OpenAI; an explicit breakpoint on Anthropic).
- D1: installations, repositories (mode: live, shadow or off), pull requests, review runs (stages, coverage, usage), findings (stable fingerprint IDs, state, GitHub comment link, feedback), users and roles, settings, audit log.
- R2: per-run context bundle, stream results, rendered review body and inline comments (so shadow runs can be inspected in the dashboard).
Trust boundaries
Section titled “Trust boundaries”- PR content, comments, tickets and tool output are evidence, never instructions. Every prompt says so, and policy comes from the base branch.
- GitHub write credentials never leave the Worker. The git container only ever sees short-lived read-only tokens.
INCOMPLETEnever approves and never dismisses an earlier block. A failed required stage is visible in the review’s coverage table.- Dashboard access requires signing in with GitHub (the CerbieAI app’s user authorization) as an active member of an allowed organisation (
CERBIE_ALLOWED_GITHUB_ORGS). Membership is checked through the app installation, matched on GitHub user id, at sign-in and hourly after (a failed check backs off for five minutes; after 24 hours without confirmation, requests are refused). Sessions are HttpOnly, SameSite=Lax cookies (only a hash is stored in D1), and API writes must come from the dashboard’s origin. Admin actions are audited.