Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

1,076 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

oh-my-bmad — a self-hosted autonomous-development platform for one operator

Telegram and a console drive a Claude Code worker through a typed, event-sourced spine — so a single person can run an agent loop they can trust, observe, and recover.

Python 3.12 uv workspace FastAPI aiogram v3 MCP mypy strict License: MIT Phase 28 closed


What this is

A platform that turns Telegram and a local console into the control surfaces for a Claude Code worker. Commands you send become typed events on an append-only log. A single-writer materializer turns the log into queryable SQLite state. Capability tiers gate every privileged action; operator approval events gate the high-risk ones. Crash the process and it rebuilds itself from the log on next start.

It's deliberately boring infrastructure — Python 3.12, FastAPI, aiogram, SQLite WAL, Docker Compose, stdio MCP. The novelty is in how the boring pieces compose, not in any one piece.

Current repo state: Phase 28 closed — Epic 107 completed the Snapshot Creation authorization runtime boundary and final closure after remote CI run 28195545005 passed. Latest tagged release is v1.3.0; this checkout contains later BMad work through Phase 28. See docs/index.md and the derivative docs/feature-status.md matrix for the current implemented/deferred feature view.

How it works (at a glance)

flowchart LR
    subgraph operator [Operator surfaces]
        TG[Telegram bot]
        CLI[console CLI]
    end

    subgraph api [HTTP API]
        RA[registry-api<br/>FastAPI · /v1/tasks]
    end

    subgraph spine [Event spine]
        LOG[(append-only JSONL log<br/>byte-stable canonical JSON)]
        STATE[registry-state<br/>SINGLE WRITER<br/>materializer + idempotency cache]
        DB[(SQLite WAL<br/>tasks · sessions · events<br/>idempotency · snapshots)]
    end

    subgraph mcp [MCP servers · stdio · capability-tier gated]
        TR[task-registry]
        SR[session-registry]
        CB[clawhip-bridge<br/>sole emission surface]
    end

    subgraph worker [Worker plane]
        WW[worker-wrapper<br/>Claude Code CLI subprocess]
    end

    TG --> RA
    CLI --> RA
    RA --> LOG
    LOG --> STATE --> DB
    LOG -.read-only tail.-> TG
    LOG -.read-only tail.-> CLI
    WW --> CB
    WW --> TR
    WW --> SR
    TR & SR & CB --> LOG
    DB --> RA
Loading

Three properties hold the whole thing together: only one writer, only-ever-append, and the bytes on disk are byte-stable. The rest of the rules in _bmad-output/project-context.md exist to protect them.

Engineering highlights

  • Event-sourced spine, byte-stable canonical JSON. Two identical envelopes serialize to byte-identical output. Replay-determinism is a property of the encoder, not a hopeful claim. (deep-dive →)
  • Single-writer invariant (FR26), statically enforced. services.<A> cannot import services.<B>. EventLogWriter is the only opener-for-write on the log. A CI gate (scripts/checks/check_imports.py) rejects PRs that drift this boundary.
  • Idempotency by UUIDv7. Every command threads the triggering UUIDv7 through a per-key asyncio.Lock → cachetools TTLCache → SQLite write-through with a 7-day retention contract. 100 concurrent retries for the same key invoke the factory exactly once. (deep-dive →)
  • Crash-injection tested. A real Docker stack gets shot at deterministic emission points; recovery is asserted to produce byte-for-byte equivalent state. Partial writes are detected and rejected by a poison-pill mechanism in the writer. (deep-dive →)
  • Capability tiers with mandatory deny-path tests. Four tiers (read / bounded-write / repo-mutation / high-risk-with-approval). Three tests mandatory per MCP tool boundary: deny-path, default-deny, escalation. @pytest.mark.security is non-skippable. (deep-dive →)
  • Multi-runtime workers. Claude Code, Codex, and Gemini adapters behind a RuntimeAdapter protocol with per-task selection, fallback, and handoff. Worker pool auto-scaling (FC-P6-1).
  • MCP tooling fleet — 9 servers. task-registry, session-registry, clawhip-bridge, git, github, verification, memory/wiki, artifact, and browser — all behind tier-gated build_server, per-server env isolation, and scoped credentials.
  • HMAC approval signing with offline verification. Operator decisions are HMAC-signed at emission; just verify-approval works offline. Key rotation with fingerprint tracking and /key-status surface.
  • Supply chain hardening. SLSA L2 provenance, cosign keyless signing, CycloneDX SBOM, license-compatibility publish gate, just verify-images — all wired into the release pipeline.
  • Metrics subscriber with cardinality discipline. Tail-loop subscriber exposes /metrics (FastAPI); 51 canonical timeseries baseline, cardinality-bounded regression test at 10K tasks, p95 <1ms.
  • mypy --strict everywhere, ruff for lint and format. No half-on rule families. Per-file ignores live in ruff.toml, not sprinkled. Bandit-S rules gate the obvious vulnerability classes (eval, pickle, yaml.load, subprocess shell=True, weak hashes).
  • AI-agent rule digest as injected context. _bmad-output/project-context.md — 386 rules across 7 categories, hand-built with multi-agent review to capture the load-bearing constraints that aren't obvious from the code alone.
  • Upstream forks behind adapter shims. upstream/omc + upstream/clawhip vendored at pinned SHAs; direct imports of vendored internals are rejected by static analysis. Contract tests gate semantic drift.
  • Three-layer secret hygiene. Pre-commit scanner + structlog sanitizer wired before the renderer + secret.accessed audit events. F-string interpolation of tokens / request bodies / PII is a banned anti-pattern in code review. Per-server env scoping ensures each MCP child sees only its own vars.

Tech stack

Layer Choice
Runtime Python 3.12 · Node.js (only inside the Claude Code worker subprocess)
Build / workspace uv ≥ 0.5 (workspace, 24 members) · just ≥ 1.14 (operator recipes)
HTTP API FastAPI (only on registry-api)
Telegram aiogram v3 (webhook + outer-middleware allowlist, ADR-0001)
Console typer CLI with full command parity
Storage SQLite + WAL · aiosqlite · Alembic (additive-only within a major)
Event log append-only JSONL, canonical JSON, fdatasync
MCP stdio by default, Streamable HTTP where configured · 9 servers (task-registry, session-registry, clawhip-bridge, git, github, verification, memory, artifact, browser)
Worker Multi-runtime: Claude Code, Codex, Gemini — supervised, auto-scaled
Observability /metrics endpoint · trace_id propagation · structlog (JSON) + sanitizer
Tests pytest · pytest-asyncio (strict) · hypothesis · crash-injection · mutation gate (cosmic-ray, 82%+) · 3200+ tests
Tooling ruff (E F I UP B SIM N + S) · mypy --strict · pre-commit · pytest-randomly
Deploy Docker Engine ≥ 24 · Docker Compose v2.24+ · SLSA L2 + cosign keyless signing
DR Litestream WAL replication to S3/B2/R2/MinIO (opt-in)

Exact versions live in uv.lock. Secrets are operator-provisioned via .env with per-server env isolation and scoped credentials.

Quickstart

# Prereqs: Docker Engine ≥ 24, Docker Compose v2.24+, uv ≥ 0.5, just ≥ 1.14
#   brew install uv just                              # macOS
#   curl -LsSf https://astral.sh/uv/install.sh | sh   # Linux (uv)

git clone https://github.com/salacoste/oh-my-bmad && cd oh-my-bmad
uv sync --frozen --all-packages    # NOT --no-dev — strips test deps
uv run pre-commit install
just bootstrap-verify              # workspace imports must be green
cp .env.example .env
$EDITOR .env                       # secrets + tunnel choice
just dev                           # macOS overlay; Linux base compose
docker compose ps                  # core services should be Up/healthy within 60s

Full deployment guides: VPS (Linux) · macOS host · Deployment entry point

How this project gets built — the BMad workflow

This codebase was produced by — and continues to follow — the BMad structured development workflow: a 4-phase lifecycle (analysis → planning → solutioning → implementation) where each phase has explicit inputs, outputs, and gates. Every artifact in _bmad-output/ is a real document produced by a real skill in that workflow.

Phase 1 — Analysis           → product-brief / PRFAQ + research
Phase 2 — Planning           → PRD + UX design
Phase 3 — Solutioning        → architecture + epics & stories + test framework + CI
Phase 4 — Implementation     → sprint plan → (create-story → validate → atdd →
                                              dev-story → code-review → trace → nfr) ×N
                              → retrospective at every epic boundary

Phase 1 took 10 epics / 88 stories, with retrospective + deferred-work governance at every epic boundary. The current repo has progressed to Phase 28 closed — event spine, multi-runtime workers, a 9-server MCP fleet, browser automation, supply-chain hardening, remote MCP transport, mTLS, historical replay, event-log lifecycle management, lifecycle-operation safety, docs/backlog reconciliation, archive-aware task history, destructive lifecycle apply readiness/product-scope planning, the read-only dashboard live-read readiness/runtime boundary series through lifecycle/snapshot, and the exact JWT-authenticated snapshot creation authorization boundary. Epic 107 is now closed by the final Phase 28 validation story. The full per-phase walkthrough, skill catalog, and "how a new feature enters the workflow" decision tree is documented separately:

➡️ docs/bmad-workflow.md — the complete workflow this project follows.

If you're an AI agent picking up new work on this codebase, that file is the process companion to _bmad-output/project-context.md (the rules digest) — read both before writing code.

Documentation map

This repo documents itself in three layers, by audience.

For AI agents (read first if you're an agent)

For humans

Deep-dives on load-bearing concepts

Planning artifacts (for the why)

What's interesting about this codebase

A few things worth a look even if you don't intend to run it:

  • The AI-agent rule digest (_bmad-output/project-context.md). A 7-category, 386-rule reference designed to be injected into the context window of a coding agent. Built collaboratively with multi-perspective review; explicitly distinguishes enforced rules (CI gates) from discipline rules.
  • The deep-dive explanations (docs/explanations/). Each one walks a load-bearing concept (event spine, idempotency, recovery, capability tiers) end-to-end with Mermaid diagrams and source-grounded code samples.
  • The single-writer enforcement (scripts/checks/check_imports.py). Service separability isn't a doc rule; it's a CI gate that statically rejects PRs that drift the boundary.
  • The crash-injection test tree (tests/crash-injection/). Real Docker stack, deterministic kill points, byte-for-byte state equivalence as the assertion.

Built with

uv FastAPI aiogram v3 Pydantic v2 SQLAlchemy SQLite structlog Hypothesis ruff mypy Claude Code

License

MIT. Use freely; attribution welcome; no warranty.

Status

Current development state — Phase 28 closed / Epic 107 done. Latest tagged release: v1.3.0. Final closure cites remote CI run 28195545005 as the Phase 28 / Epic 107 evidence gate. The canonical status source is _bmad-output/implementation-artifacts/sprint-status.yaml; docs/feature-status.md is the derivative human-readable implemented/partial/deferred matrix.

Phase Scope Epics Current status
1 Core platform — event spine, registry, Telegram, console, workers 1–7.5 Done
2 Observability — supply chain, trace_id, metrics, HMAC approvals, budget enforcement, Litestream DR 8–13 Done
3 MCP tooling fleet — git, github, verification, memory/wiki, artifact 14–19 Done
4 Browser automation — Playwright MCP, screenshot, tab management 20–22 Done
5 Multi-runtime — Codex, Gemini adapters, per-task selection, handoff 26–29 Done
6 Server execution pool — Postgres, state machine, multi-worker 30–34 Done
7 Reliability hardening — heartbeat detection, structured output, env isolation 35–40 Done
8 Platform hardening & debt closure — zero open deferred items 41–45 Done
9 Operational excellence — PR drafts, runbooks, stale TODO cleanup 46–48 Done
10 Remote MCP transport — Streamable HTTP + bearer auth 50–55 Done
11 mTLS — internal Docker-network TLS profile and CA tooling 56–59 Done
12 Historical event replay — replay engine, validation, snapshots, task history 60–63 Done
13 Event log lifecycle — archive manifest, hot+archive replay, package streaming 64–68 Done
14 Event log lifecycle operations — ADR-0025, non-destructive dry-run, hot-only task-history boundary 69–73 Done
15 Lifecycle documentation reconciliation and backlog triage 74–75 Done
16 Archive-aware task history — read-only hot+archive history query 76–80 Done
17 Destructive lifecycle apply readiness — planning/safety contract only, no destructive implementation 81–85 Done
18 Destructive lifecycle apply product scope — PRD/status-only gate plus next non-destructive candidate selection 86 Done
19 Read-only dashboard shell — static dashboard panels, read-only/no-mutation boundaries 88–92 Done
20 Dashboard live-read contracts — route metadata, provenance/freshness states, unavailable aggregate/session decision 93–97 Done
21 Dashboard rendering readiness — view models, fixture/static rendering, live-read wiring decision gate 98–100 Done
22 Health/readiness runtime boundary — narrow GET /v1/health browser runtime 101 Done
23 Task-detail runtime boundary — narrow GET /v1/tasks/{task_id} browser runtime 102 Done: 102.1 route selection, 102.2 runtime boundary, and 102.3 final closure
24 Phase 24 — Event timeline / transitions runtime boundary for exact GET /v1/tasks/{task_id}/events and GET /v1/tasks/{task_id}/transitions 103 Done
25 Phase 25 — Trace correlation runtime boundary for exact GET /v1/trace/{trace_id} 104 Done
26 Phase 26 — History / Replay runtime boundary for exact GET /v1/tasks/{task_id}/history, GET /v1/events/replay, and GET /v1/events/replay/validate with visible replay target query discipline 105 Done
27 Phase 27 — Lifecycle / Snapshot runtime boundary for exact GET /v1/events/replay/snapshots plus passive lifecycle-readiness evidence display 106 Done: Story 106.3 records final closure with remote CI run 28139358221
28 Phase 28 — Snapshot Creation authorization runtime boundary for exact JWT-authenticated POST /v1/events/replay/snapshots 107 Done: Story 107.3 records final closure with remote CI run 28195545005

Destructive lifecycle apply is still unimplemented, and object-storage lifecycle jobs plus scheduled retention remain future work. Dashboard runtime wiring remains intentionally narrow: health/readiness, task detail, event timeline/transitions, trace correlation, history/replay, lifecycle/snapshot listing, and the single visible JWT-authenticated snapshot-create affordance only. Aggregate/session/digest, task-list/search/discovery, broad dashboard wiring, destructive lifecycle mutation, archive/manifest mutation, snapshot deletion/restore, and any other mutation/control affordances remain deferred/fail-closed unless a later BMad phase explicitly approves them.

Issues and discussion welcome — security reports per SECURITY.md.

About

Self-hosted autonomous-development BMAD framework platform — Telegram + console drive Claude Code through a typed, event-sourced spine. Python 3.12 · FastAPI · aiogram v3 · MCP stdio.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages