Skip to content

Architecture

Cerbie runs entirely on Cloudflare. The goal is the best review in the shortest wall-clock time. Every design choice below either cuts latency (parallelism, warm mirrors, no installs, prompt caching) or raises signal (specialists, independent verification, convergence across rounds).

flowchart TD
  GH[GitHub webhooks] --> W[Worker: verify + gate]
  W --> DO[PullRequestCoordinator<br/>one Durable Object per PR]
  DO -->|debounce, supersede| WF[ReviewWorkflow]
  WF --> CTX[Context step]
  CTX --> S1[Summary]
  CTX --> S2[Primary reviewer<br/>sharded for large PRs]
  CTX --> S3[Specialists:<br/>security · data · contracts]
  CTX --> S4[Custom checks]
  CTX --> S5[Deterministic:<br/>deps · lockfile · secrets]
  CTX --> S6[Follow-up of<br/>earlier findings]
  S2 & S3 & S4 & S5 & S6 --> V[Verifier<br/>fresh context]
  V -->|risky PR that would be approved| C[Approval challenge<br/>+ verification of its findings]
  V --> P[Publish: review, inline comments,<br/>check runs, thread housekeeping]
  C --> P
  S1 -.->|posted as soon as ready| GH2[Sticky summary comment]
  S2 & S3 & V -.->|tools| GIT[(Git mirror container<br/>shallow bare clones)]
  S2 & S3 & V -.->|LLM calls| LLM[cerbie-llm-gateway<br/>CLIProxyAPI container]
  P --> D1[(D1: runs, findings,<br/>users, settings)]
  WF --> R2[(R2: context, results,<br/>rendered output)]
  1. Webhook (apps/worker/src/index.ts, webhooks.ts). The Worker verifies the HMAC signature and deduplicates by delivery ID in D1. It then reads .cerbie/config.yaml from the base branch to decide whether this event should trigger a review (drafts, ignored authors and labels, base branches, cerbie: skip). Commands (@cerbieai …) and thread replies are handled here too.
  2. Coordinator (coordinator.ts). One Durable Object per PR debounces bursts of pushes using a single alarm. If a newer head arrives while a review is running, the old Workflow is terminated (no tokens spent on stale code) and the new head is reviewed. Re-delivered or duplicate events for an already-reviewed head are ignored.
  3. Workflow (review/workflow.ts). A Cloudflare Workflow makes every stage durable and independently retried:
    • start: creates the Cerbie (Review) check run (in progress) and, on a PR’s first review, a placeholder summary comment.
    • context: fetches PR metadata, files and patches, CI status, discussion, the linked Linear ticket, guideline files, .cerbie/, and the PR’s earlier findings. In parallel it syncs the git mirror to the exact base and head commits. The bundle goes to R2. Triage (deterministic, zero latency) picks the tier and streams.
    • Fan-out, all in parallel: summary, reviewer streams, custom checks, deterministic checks, follow-up of earlier findings. The summary is posted the moment it is ready.
    • verify: candidates are merged and deduplicated. Duplicates of still-open earlier findings are dropped so threads aren’t repeated. A verifier with a fresh context confirms, adjusts or rejects every critical and should-fix candidate against the code. The confidence floor and suggestion cap apply.
    • challenge: only when the PR is risky (deep tier, or any specialist ran) and the verified findings would still approve it. One reviewer gets the first pass’s notes, files read and findings, and looks for failure modes nobody traced. Its should-fix and critical findings go through a second verification (verify challenge). If the challenge fails, coverage is incomplete and the PR is not approved.
    • publish: the head guard skips publication if the PR moved. Cerbie then posts one review (verdict, inline comments with suggestions) and completes the check runs (one per custom check). It leaves earlier inline comments untouched and never resolves threads (that is the PR author’s call); for a fixed finding whose thread someone replied to, it posts a short “✅ Addressed in …” reply. It also dismisses stale Cerbie blocks and updates the summary with the verdict.
  4. Replies (reply-workflow.ts). A human reply on a Cerbie thread, or @cerbieai <question>, runs a short agent with the same tools. The agent re-reviews the finding(s) a reply concerns with the reply as context: it may withdraw, confirm a fix, accept a reasoned deferral (should-fix and suggestions only) or uphold the finding, and one top-level comment can cover several findings. It is capped at three replies per thread. Its decisions go to the PR’s coordinator, which applies them (applyDecisions) and re-decides the verdict (reassessVerdict) in one serialised step, never concurrently with a review; if nothing blocks, it approves without re-reviewing the PR. Withdrawn and deferred findings are not raised again unless they come back more severe. Hand-resolved threads do not change findings.

The Cerbie CLI runs the same ReviewWorkflow before a pull request exists:

  1. cerbie login uses GitHub device flow through the CerbieAI app (/cli/v1/login/*). The Worker keeps only a hash of the CLI session token, and checks every request against the user’s permission on the repository, matched on their immutable GitHub ID.
  2. cerbie review uploads a thin git bundle (merge base..head; uncommitted work becomes a temporary commit) to /cli/v1/repos/:owner/:name/reviews. The Worker reserves the user’s rate-limit quota, stores the bundle in R2 and starts the workflow with params.local (PR number 0, never published).
  3. The context step syncs the merge base from GitHub, imports the bundle into the git mirror and diffs there. Policy comes from the tip of the target branch on GitHub, as for a PR. Git objects are content-addressed, so the reviewed commit is exactly the code its SHA names.

Follow-up reviews are incremental (files changed since the last reviewed commit) unless decideReviewScope (packages/core/src/scope.ts) says otherwise. It uses signals that need no model call: merge base, a hash of the repository’s review rules, whole-PR triage, delta share, rounds and age since the last full review, and relative importers of changed exports, read from the git mirror. Each run’s context records the rules hash and whole-PR triage for the next decision; eval/scope-replay-2026-10-08.md replays production history through it.

Every run records a cache key: merge base, head tree and a policy hash (.cerbie/ and guideline files, org instructions, the linked ticket as read, recipe, prompts). A PR’s first review (or any local review) that finds a complete, non-incremental run with the same key from the last 14 days publishes that run’s findings instead of reviewing again; only the summary is written fresh. If an older run of the same code is still in progress, the workflow sleeps and waits for it. Runs only wait for older runs, so two can never wait on each other.

Tier When Streams
light Docs, tests or lockfiles only, or ≤60 substantive changed lines with no risk signals One focused reviewer
standard Everything else Primary reviewer, plus specialists whose signals fire
deep Configured paths.risk match, >600 substantive lines, or both security and data signals Primary (sharded across up to four reviewers for big diffs) plus every signalled specialist

Signals come from paths (auth, migration, routers/, …) and added code ($transaction, dangerouslySetInnerHTML, exec(, removed exports, …). Triage errs towards running a specialist: a parallel call that finds nothing costs far less than a missed lens. The deterministic checks, custom checks and follow-up always run when relevant.

Installing dependencies for the monorepo takes minutes and gigabytes, and CI already runs type checks and tests. Cerbie reads CI status instead and spends its budget on what CI can’t catch: logic, authorisation, data integrity, concurrency and contracts. Agents get a read-only toolset over exact commits:

Tool Backed by
read_file, list_directory git mirror (git show/ls-tree); falls back to the GitHub contents API
grep, find_files git grep/ls-tree against the commit object, with no checkout (full monorepo grep takes about 0.5s)
diff_file, git_log PR patches or git diff; shallow history

The git mirror (containers/git) is a small Bun server. It keeps one shallow, bare clone per repository in a container that stays warm for 30 minutes. Cold sync of the monorepo takes about 5s, a warm incremental sync about 10ms. It receives a per-request, repository-scoped, read-only installation token and never executes repository code.

Every LLM call goes through cerbie-llm-gateway, a CLIProxyAPI instance in a Cloudflare Container reached over a service binding. Model assignments are a recipe (apps/worker/src/llm.ts). Admins can override it in the dashboard without a deploy.

Role Default Why
Every role defaults to GPT-6.1 Sol; always prefer it over GPT-6 Sol. Roles differ in effort and step budget.
Role Default Why
primary GPT-6.1 Sol, high effort Deepest tracing of behaviour
challenge GPT-6.1 Sol, high effort Runs after verification, only on risky PRs (deep tier or any specialist) whose verified findings would still approve. It reads the first pass’s notes and findings and hunts for failure modes nobody traced. Its findings are verified separately; if it fails, the review is incomplete rather than approved
specialist GPT-6.1 Sol, high effort Focused lenses in parallel
verifier GPT-6.1 Sol, high effort Re-checks each finding in a fresh context with its own tools
light GPT-6.1 Sol, high effort Small PRs still get the strongest reviewer
summary, reply GPT-6.1 Sol, low / medium effort Speed

The large shared context (PR brief, guidance, numbered diff) is the first message in each conversation, so every tool-loop step reuses it from the prompt cache (automatic on OpenAI; an explicit breakpoint on Anthropic).

  • D1: installations, repositories (mode: live, shadow or off), pull requests, review runs (stages, coverage, usage), findings (stable fingerprint IDs, state, GitHub comment link, feedback), users and roles, settings, audit log.
  • R2: per-run context bundle, stream results, rendered review body and inline comments (so shadow runs can be inspected in the dashboard).
  • PR content, comments, tickets and tool output are evidence, never instructions. Every prompt says so, and policy comes from the base branch.
  • GitHub write credentials never leave the Worker. The git container only ever sees short-lived read-only tokens.
  • INCOMPLETE never approves and never dismisses an earlier block. A failed required stage is visible in the review’s coverage table.
  • Dashboard access requires signing in with GitHub (the CerbieAI app’s user authorization) as an active member of an allowed organisation (CERBIE_ALLOWED_GITHUB_ORGS). Membership is checked through the app installation, matched on GitHub user id, at sign-in and hourly after (a failed check backs off for five minutes; after 24 hours without confirmation, requests are refused). Sessions are HttpOnly, SameSite=Lax cookies (only a hash is stored in D1), and API writes must come from the dashboard’s origin. Admin actions are audited.