Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
70 changes: 70 additions & 0 deletions evidence/security/osv-baseline-2026-08-31.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# OSV offline baseline — 2026-08-31

This is a qualified repository-security baseline, not a product verdict and not
an allowlist. It records what OSV-Scanner observed so remediation can be
measured without hiding existing findings.

## Reproduction

- Source revision: `225480e5b6a018014bef1c190ed8ac3914a472e7`
- Scanner: OSV-Scanner `2.5.1`
- Command: `pnpm quality:vulnerabilities`
- Network during scan: disabled
- Warm scan duration: 8.942 seconds
- Machine output: ignored `artifacts/tooling/osv/results.sarif`
- Receipt: ignored `artifacts/tooling/osv/receipt.json`

The command exits `0` when clean, `1` when findings exist, and `2` for an
operational failure. Database refresh is deliberately excluded from the scan.

## Database identity

| Ecosystem | Bytes | SHA-256 |
|---|---:|---|
| crates.io | 3,415,816 | `e276fb0061eefcb63ce5ba63b6888640ccc809dc195a698af7b7dbaaeaa9a02c` |
| Go | 11,481,601 | `49b2410b903ae009b3a89ab9be5245d531d0c176b6809931f2182b5ee6bc1204` |
| npm | 221,667,534 | `d89fbb49609224a0ecdb5cd14d6b874ab105197a5e83543f78450bab406cceb8` |

## Findings

- 35 affected package versions across two lockfiles
- 52 advisory/package matches
- 49 unique primary advisory IDs in OSV JSON
- 48 unique SARIF rules after alias normalization
- SARIF severity by unique rule: 0 critical, 11 high, 18 medium, 2 low,
17 without a numeric severity
- Root `pnpm-lock.yaml` and the Go module were clean

`docs-site/pnpm-lock.yaml` contains 17 affected package versions and 33
advisory/package matches. They are transitive dependencies of the single direct
docs-site dependency, Blume `1.0.4`. Upgrading Blume is a production/build
dependency change and requires owner approval.

`apps/desktop/src-tauri/Cargo.lock` contains 18 affected package versions and 19
advisory/package matches. Most are unmaintained GTK3 bindings reached through
Tauri's Linux target graph. `event-listener` is reached through the Linux
notification stack. The `unic-*` crates are reached through
`urlpattern -> tauri-utils` and are not classified as Linux-only. These are
reachability notes, not suppressions.

## Gate decision

Do not add a second vulnerability scanner or silently baseline these findings.
First test the Blume update and bounded Cargo lockfile remediation, then repeat
this exact offline scan. CI enforcement should follow remediation so the gate
does not normalize known high-severity results.

## Approved maintenance follow-up

At exact revision `855202998b56c1658b9decda22298a1b63fb5caf`, after updating
Blume 1.0.4 to 1.5.3 and event-listener 5.4.1 to 5.4.2, the same database set
and offline runner completed in 9.745 seconds with 39 result instances and 36
normalized SARIF rules. This is a reduction from 52 advisory/package matches,
not a clean result and not an allowlist.

The independent docs-site production audit still reports 12 high, 7 moderate,
and 1 low advisory. Current Blume transitively retains those build/documentation
paths, and several have no resolution through the current direct version. The
repository-wide root production audit remains clean because the docs site has
its own lockfile; both facts must stay visible until issue #195 completes
reachability and upstream/remediation review.
45 changes: 45 additions & 0 deletions evidence/security/trivy-config-qualification-2026-08-31.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# Trivy configuration-scan qualification — 2026-08-31

## Scope

This trial asks whether Trivy's embedded misconfiguration checks add a useful
repository lane without downloading databases or overlapping the offline OSV
scanner. It does not qualify Trivy vulnerability scanning or product bundling.

## Tool and safety flags

- Tool: Trivy 0.74.0
- License: Apache-2.0
- Command lane: `trivy config`
- Explicit controls: `--disable-telemetry`, `--skip-version-check`, and
`--skip-check-update`

Debug output confirmed that the notification/version request was skipped, no
downloadable checks were loaded, and 563 embedded checks were used. This is
command-level evidence, not packet-capture proof of zero outbound traffic.

## Results

The unbounded repository scan found only a Dockerfile inside
`docs-site/node_modules/yaml-language-server` and reported two findings against
that dependency-owned file. After excluding dependency, build, target, and
artifact directories, Trivy detected one file:

`apps/desktop/tests/fixtures/warm-verification/differential-runtime-qualification-current.json`

It classified that product test fixture as CloudFormation, reported 24 passing
checks and zero failures, and found no first-party Dockerfile, Terraform,
Kubernetes, Helm, Ansible, Azure ARM, or CloudFormation surface to protect.

The SARIF 2.1.0 envelope itself is structurally complete. In JSON debug output,
Trivy also includes repository URL, branch, commit, author, and committer
metadata, which is acceptable for this public-repository trial but must be
considered before any private-code product use.

## Decision

Trialled, not wired. A permanent lane would currently scan dependency-owned
files or misclassify a qualification fixture while protecting no supported
first-party infrastructure configuration. Re-evaluate when the repository adds
a real supported IaC or container surface. Keep all three network-suppression
flags mandatory in any future trial.
60 changes: 60 additions & 0 deletions evidence/verification/apple-container-qualification-2026-08-31.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Apple Container qualification — 2026-08-31

## Scope

This receipt qualifies Apple's `container` CLI as an external-prerequisite
sandbox candidate on one supported development host. It does not add a product
dependency, bundle the runtime, or approve a release architecture.

## Identity

| Item | Observed identity |
|---|---|
| Host | macOS 27.0 build 26A5421a, arm64 |
| Installer | `container-1.3.1-installer-signed.pkg` |
| Publisher SHA-256 | `a7c1b9d7927d30875f2f6c7bd1d0cb06c2daa6ca57ce9e90a5144e898fdf54a8` |
| Signature | Developer ID Installer: Apple Inc. - Containerization (UPBK2H6LZM); trusted, notarized, timestamped |
| CLI | 1.3.1, release commit `a9a62e2` |
| API server | 1.3.1, release commit `a9a62e28f6beb88940122a3d7b286f2d5ae8053a` |
| Default kernel | Official Kata kernel 3.32.0 |
| Trial image | `alpine:3.22`, observed Alpine 3.22.5, local digest prefix `14358309a308` |

The package digest matched the release publisher value before installation.
`pkgutil --check-signature` and `spctl` both accepted the downloaded package.
Installation required the owner's administrator authorization. The first
service start separately downloaded and verified the 664.3 MB default kernel.

## Measured trial

- First image/init-image run: 20.56 seconds wall time, including registry fetch
and unpack; the kernel had already been provisioned.
- Warm cached no-op run: 0.61 seconds wall time.
- Limits exercised: 1 CPU and 256 MB memory.
- Isolation exercised: an internal network plus `--no-dns`, read-only root,
read-only bind mount, and all Linux capabilities dropped.
- The controlled network probe failed, host-only environment marker was absent,
`/Users` and `/root/.ssh` were absent, and writes to both the mounted workspace
and `/root` failed with a read-only-filesystem error.
- `--rm` teardown left zero containers and zero local volumes. The cached image
and init image occupied 1.45 GB and were intentionally retained.

The trial network was deleted after use. No host credential, home, SSH, cloud,
or production path was mounted or inspected.

## Failed requirement: caller-owned path containment

The CLI accepted a controlled bind source containing `..` when that source
resolved to an existing sibling fixture. Apple Container therefore provides
mount mechanics, but it does not enforce CodeVetter's workspace-root policy.
Any adapter must canonicalize both the allowed root and requested source,
reject sources outside the allowed root before process launch, pass the
canonical source to the CLI, and cover symlink and time-of-check/time-of-use
cases. The product must not treat the CLI's mount validation as containment.

## Decision

Keep Apple Container as the measured external-prerequisite candidate, not a
bundled dependency. Its isolation controls and warm-start result are promising,
but path containment must be owned by CodeVetter. Issue #197 remains open for
the CLI-versus-Containerization-versus-libkrun architecture comparison, idle
resource measurement, signing/notarization analysis, and an adapter design.
93 changes: 93 additions & 0 deletions evidence/verification/cross-review-benchmark-2026-09-02.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
{
"schema_version": "codevetter.cross-review-benchmark-summary/v1",
"recorded_at": "2026-09-02T15:17:50.495Z",
"source": {
"runner_revision": "017985904f89ebcb4c73fe00147e0bb4fde53def",
"engine_revision": "f812ba1c8d5d349bc9a38e2ebde9f03d54509444",
"binary_sha256": "113e0a3fa419a4487af0cec8515d9f0fc9534d906c9f0f5b6ea127cd03e518e0",
"corpus": "benchmarks/public-catch-rate",
"cases": 27,
"expected_labels": 29,
"task": "Review this exact change for correctness, security, reliability, and maintainability defects."
},
"scoring": {
"status": "human_reviewed",
"deterministic_proposal": "same source path, label line within five lines, and narrow defect-type keywords",
"human_correction": {
"case_id": "py-path-traversal",
"reviewer": "codex",
"finding": "Untrusted report name permits arbitrary local file disclosure",
"reason": "The core claim and source identify the labeled traversal defect even though the narrow mapper keyword was absent."
},
"confirmed_single_reviewer_miss": {
"case_id": "ts-dead-code",
"reviewer": "codex",
"label": "unused-helper-function",
"severity": "low"
}
},
"reviewers": {
"claude": {
"caught": 29,
"expected": 29,
"findings": 134,
"strict_precision": 0.21641791044776118,
"recall": 1.0,
"f1": 0.35582822085889565,
"total_duration_ms": 2387162,
"mean_duration_ms": 88413,
"usage_available_cases": 0
},
"codex": {
"caught": 28,
"expected": 29,
"findings": 46,
"strict_precision": 0.6086956521739131,
"recall": 0.9655172413793104,
"f1": 0.7466666666666666,
"total_duration_ms": 2675832,
"mean_duration_ms": 99105,
"usage_available_cases": 0
},
"cross": {
"caught": 29,
"expected": 29,
"findings": 99,
"strict_precision": 0.29292929292929293,
"recall": 1.0,
"f1": 0.45312500000000006,
"total_duration_ms": 5063091,
"mean_duration_ms": 187522,
"usage_available_cases": 0,
"finding_classes": {
"corroborated": 16,
"claude_only": 63,
"codex_only": 1,
"conflicting": 19,
"rejected": 0,
"stale": 0,
"unresolved": 0
}
}
},
"execution": {
"completed_cases": 27,
"scored_incomplete_receipts": 0,
"transient_incomplete_attempts": 1,
"retry_case": "py-insecure-deserialization",
"raw_artifact_bytes_approximate": 1153434,
"temporary_repositories_retained": false
},
"decision": {
"default_strategy": "claude_unchanged_for_compatibility",
"efficiency_candidate": "codex",
"cross_review": "optional_high_recall",
"reason": "Cross-review recovered one low-severity label over Codex but added 53 findings and about 88 seconds per case."
},
"limitations": [
"The corpus is synthetic and single-file; it does not establish real-repository or real-PR quality.",
"Strict defect-only precision penalizes process findings and plausible additional defects outside the narrow ground truth.",
"Provider usage was absent from every pass summary, so observed cost is unavailable.",
"Reviewer agreement is review coverage and never executable proof."
]
}
58 changes: 58 additions & 0 deletions evidence/verification/native-agent-island-settings-2026-09-02.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Native Agent Island settings parity — 2026-09-02

## Verdict

The Agent Island configuration contract is transferred. Rust, the
`codevetter settings` CLI, the native Settings desk, and the retained supervised
helper use the same 12 non-secret preference keys, defaults, and bounded option
sets. Agent Island remains opt-in and off by default.

The live runtime is not transferred. The new Evidence Workbench does not launch
the helper, inspect live session content, speak updates, preview real provider
output, or action provider requests. Agent and MCP surfaces have no Agent Island
authority. The native desk therefore labels configuration as live and runtime
transfer as pending.

## Canonical preference contract

| Key | Default | Allowed values |
| --- | --- | --- |
| `native_agent_island_enabled` | `false` | boolean |
| `native_agent_island_speech_muted` | `false` | boolean |
| `native_agent_island_speak_completion` | `true` | boolean |
| `native_agent_island_speak_attention` | `true` | boolean |
| `native_agent_island_speak_failure` | `true` | boolean |
| `native_agent_island_speech_volume` | `0.8` | `0.5`, `0.8`, `1` |
| `native_agent_island_speech_rate` | `0.48` | `0.4`, `0.48`, `0.56` |
| `native_agent_island_speech_cooldown` | `30` | `15`, `30`, `60` seconds |
| `native_agent_island_quiet_start` | off | off, `20`, `21`, `22`, `23` |
| `native_agent_island_quiet_end` | off | off, `6`, `7`, `8`, `9` |
| `native_agent_island_codex_voice` | empty | at most 256 non-control characters |
| `native_agent_island_claude_voice` | empty | at most 256 non-control characters |

The receipt schema remains `codevetter.native-settings/v1`. Unknown keys,
invalid options, duplicate keys, partial Agent Island contracts, and projected
`github_token` values fail closed. No session, prompt, output, command, path,
provider response, credential, or voice sample enters the preview.

## Qualification

- An isolated-app-data CLI smoke listed exactly 12 `agent_island` rows, saved
only `native_agent_island_enabled=true`, returned that exact `saved_key`, and
continued to declare `github_token` excluded.
- Rust unit coverage verifies the 12-row contract, save round trip, and the
256-character voice-identifier bound.
- `pnpm test:native` passed 76 Swift package tests with no failures or skips in
29.8 seconds; the native Debug app then compiled in 2.2 seconds.
- Swift rejects a schema-v1 receipt missing even one Agent Island preference.
- Deterministic 2560x1600 dark and light renders were inspected. The light pass
found and fixed inherited dark text inside the always-black status capsule.

| Render | SHA-256 |
| --- | --- |
| `settings-agent-island.png` | `df36cb66ed120f7eac2e0387379484e54fdf67e99f1a68e0daa5a820227bf158` |
| `settings-agent-island-light.png` | `2073326d72cfd6fee7a1ed64074663388f2f1381d11c53ee8a5610c4398e0799` |

No helper, installed application, foreground automation, provider process,
network listener, release, signing, notarization, or deployment action was
started by this qualification.
54 changes: 54 additions & 0 deletions evidence/verification/native-application-service-2026-09-01.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# Native verification application-service qualification

Date: 2026-09-01
Scope: Review plan, execute, progress, cancellation, and terminal receipt transport

## Result

Native Review and `codevetter check` now enter one Tauri-independent Rust
application service instead of independently assembling the verification
lifecycle. One bounded request id correlates:

- `codevetter.verification-command/v1` input;
- monotonic `codevetter.progress/v2` events;
- request-scoped `codevetter.verification-cancel/v1` termination; and
- the distinct canonical preflight or final local-check receipt.

The Rust service still delegates source identity, target discovery, execution,
persistence, verdicts, and limitations to the existing authoritative engines.
Swift owns process supervision and rendering only.

## Executable proof

- a real clean-clone `codevetter check --preflight --request-id
native-service-live-smoke --json` returned `ready`, preserved the exact
request id, and resolved immutable base/head SHAs;
- the same command against the dirty migration worktree failed closed before
producing a receipt;
- Rust service tests cover bounded/generated request ids, progress ordering,
cancellation identity, and a real two-commit Git preflight through the
service;
- all 28 CLI tests cover parsing, progress-v2 serialization, shared-fixture
receipt/exit parity, and existing
output/exit semantics;
- Swift tests prove exact CLI arguments, matching progress and receipt
decoding, foreign-progress rejection, foreign-cancellation refusal,
matching cancellation without a receipt, mismatched-receipt rejection,
1,000-event throughput, and worker crash recovery;
- the final native background gate passed 61 Swift tests and the macOS Debug
application build;
- the final all-target Rust regression passed 1,077 tests with 31 intentional
ignores and no failures; and
- the fresh Release host and package qualifier passed at
`artifacts/native-package/qualification-XrpqmY/CodeVetter.app`.

No test failure is waived by this receipt.

## Authority boundary

MCP remains read-only and does not gain Review execution or cancellation
authority. It can now retrieve one already-persisted canonical local-check
receipt by bounded run id inside its authorized repository scope. Cancellation
is a supervised transport terminal action rather than a persisted engine
event. This contract does not imply a daemon, concurrent multi-run scheduler,
release authorization, or Tauri retirement.
Loading
Loading