Claude Evolution System
Self-improving AI development environment. Discovers new capabilities, evaluates them against a scoring framework, and integrates approved tools on cron.
- Autonomous capability discovery pipeline
- 5-criterion evaluation scoring framework
- Auto-integration of approved capabilities
- Multi-model orchestration (Claude, Codex, Gemini)
Activity Timeline
35 agent definitions catalogued; workspace codex-researcher found to shadow global version (critical collision). 15 rationalization findings (2 P1, 6 P2, 7 P3). Weekly-insights ran zero novel findings — duplicate digest from 3 prior runs.
LFM25 registry reflects latest open-weight entries. Self-observation UI (1,116 lines) completed with post-gate bugfixes. Sol-cutover remediation ongoing. 55 evals pending.
Evolution pipeline degraded but stable. Backlog at normal steady-state size.
Burn-guard sessions removed 4 stale facts across 19 skills. No new GA model releases. 55 pending evaluations tracked across the system.
GPT-5.6 family and Gemini 3.6 Flash GA dates and pricing confirmed against live vendor docs. Discord monitoring and Brave/Exa/Pro escalation routing specified as an Investigation Runner agent framework — not yet deployed.
Daily capability discovery heartbeat job templated with Phase 1 sources defined. GPT Max synthesis bypass fix confirmed with 22/0 contract tests. Investigation Runner and Herald agent specs drafted for Discord topic processing and backlog surfacing.
Capability discovery pipeline has 54 items waiting without processing across 5 active sessions. No commits. Agent roster at 27 active with 71 registered skills.
M7 completion-stamp mechanism tracks report status without production ledger writes. OAuth-orphaned reports cleared. Night-shift-sentinel alert classifier updated. Templated renderer copy discovery revises distinct session load estimate to ~130/day.
61 KB report produced. tmux 3.4 libsystemd scopes give verified 1:1 pane identity (36/36). Root cause: identity derived post-acquisition rather than at mint-time. Zombie session bg-c-bq-041 — 42 hours past TTL, 18 orphaned processes — cited as primary evidence.
Write-hook fail-closed, preflight refusal honesty, and private-checkout wording tightened. Two P1 findings (write-hook fail-open, owner-interest gate) and four additional P2s under active remediation; 52 evaluations pending.
Autonomous executor run logged. Five ordered burn tasks initiated including grok-build flag removal and commission-orphan requeue. Infrastructure sitting at 51 pending evaluations, 27 agents, 71 skills on Claude Code v2.1.228.
30-cell Fable ladder executed and manually audited, converging on 10 confirmed defects. Reasoning-effort research comparing Fable, Opus 5, and Sol 5.6 commissioned on inverse-scaling patterns. Evaluation backlog steady at 51 pending items.
v2.1.226 current, 50 evaluations pending. Sweep-card now produces narrative output rather than structured dumps. Two investigations queued: queue-burn drain-rate analysis and row-343 deep-dive with concrete options.
Commit 95edb13 records autonomous orchestrator activity for the day. Version audit covers 2.1.217–2.1.224 range since baseline 2.1.216.
Hybrid GPT-5.6 Luna approach activated for closure-verify and silence-sweep paths; all misclassifications conservative-direction. False-claim audit corrected post 067. Dispatch bug identified: 18 of 146 closed items remain dispatchable.
Defect hunt identified liveness alarm stuck 10 days, session reaper with empty nominations list running nightly, and investigation runner burning ~$155/month on silent channel since July 24. Root cause: hand-maintained queue registry. Interview-mode design accepted and BUILD authorized.
Autonomous executor run committed (fd02004). Pipeline holds 49 pending evaluations across 27 agents and 71 skills — expected standing load.
Gates A and B completed; C (exit 141) and D (exit 3) need triage. Scout cost-gate reworked to per-task AA intelligence index priority. Repository reseat approved for three repos; 49 evaluations pending.
Model-scout dry run processed 414 catalog records and ranked 32 candidates; cost gate refactored to per-task comparison with fail-closed enforcement. ask-polaris prototype enables background agents to route blocking decisions to orchestrator without freezing the session.
New tool: 4 configs, 6 modules, 134 tests, 414-record dry run in ~4s. Live run caught Inkling-Small throughput shortfall, DeepSeek V4-Flash availability gap, and credential scope escalation. CONTEMPORARY-MODELS-CHECK passed with zero table updates needed.
Fusion embedding approach degrades retrieval quality — verdict: do not adopt. Shadow adjudicator accuracy drops from 67.3% to 61.2% when the governance constitution is loaded; three-step baseline-first procedure now documented to correct the conservative bias.
146 of 235 backlog rows triaged; ~37 gated on owner decisions, not model capability. 71 of 91 cross-session asks were pure repetition from session-boundary context loss. Multi-arm benchmark (Haiku/Fable/Sol-single) run on frozen-corpus snapshots with regression controls.
Double-counting in synonym pairs corrected in term-filtering lens. Stamp-field markdown misparse fixed in backfill reopen logic. Owner-interest routing gate reopened 10 archived rejects. 146+ backlog rows triaged into burn (48) and codex-ultra (12) queues.
R1–R15 routine items plus N2/N4/S3/S2 committed to integration branch. N1 (4 posts) held for content-court review. Contemporary model check confirmed GPT-5.6 Sol/Terra/Luna and Gemini 3.6 Flash all current.
Identity-linkage terms, staging IPs, and Moltbook registration credentials purged from mirror source via 23 substitutions. Plaintext API key recovery triggered the audit. Executor run completed cleanly.
All 5 background tasks exited cleanly. Housekeeping approvals unlocked the integration/reattach-approved-backlog branch. Roster: 27 agents, 69 skills, 50 pending evaluations queued.
Pipeline state file lagged global config on Gemini model version — discovery filed. Night-shift TOCTOU race and four additional defects fixed with probe validation and passing test suites. Discord investigation runner architecture specified.
Model confirmed against Google docs July 21: 1M context, $1.50/$7.50/M. Daily heartbeat capability-discovery job configured. Polaris-C (Fable) launched on Account B with full authority stack inherited.
Investigation runner fetches Discord channel messages, filters topics, dispatches research, writes structured markdown reports. Daily integration executor armed on codex-exec transport. Model reference table validated against current GA dates.
50 backlog items processed, 16 techniques added to library index across four sessions. Eight commits covered 2.1.193/195 registry updates, hook-matcher audit, Capframe, and defense-in-depth notes. Gemini 3.5 Pro detected in freshness sweep and queued for evaluation.
15-version gap (v2.1.200–v2.1.214) closed in one session. 9 discovery files generated for pipeline evaluation. 46 evaluations remain in the pending queue.
State file current as of 2026-07-16. 46 pending evaluations, 25 agents, 64 skills active.
Fable 5 added, Opus 4.8 and GPT-5.6 Sol/Terra/Luna registered. Discovery file created for agent/skill references needing model ID updates. Evaluation backlog and pending discovery sweep keep status degraded.
Design day: 8 sessions, 0 commits. Two Discord-integrated agents specified. Decision-making pattern established (AskUserQuestion for all forks). Evaluation loop stalled at 40 pending items.
Hardened approval gate with unified orchestrator safety model. Added backlog test-and-deploy method with Opus rubric and preregistration tooling. 40 items queued for next evaluation cycle.
Model landscape updated with GPT-5.6 variants announced June 26. Agent specs drafted: Investigation Runner for Discord inbox scanning, workspace orchestrator with 5-tier priority matrix. 43 evaluations pending in queue.
Phase 1 of publish model inversion fully designed but held pending owner approval. Portfolio cleanup shipped in 3 commits: README, private mirror backup, registry path and IP fixes. Model version state file found 6 weeks stale relative to CLAUDE.md reference.
ANTHROPIC_BASE_URL export during RC launch was causing proxy 404s on availability checks — stripped from launch script. Allow-list publishing model ready for Phase 2 review. Multi-account isolation verified via CLAUDE_CONFIG_DIR workaround.
RC was blocked by ANTHROPIC_BASE_URL proxy in launch script — patched. Newspaper crontab failed due to missing logs directory — fixed, job 4 completed. ChatGPT Pro MCP failure modes characterized (timeouts, CDP drops, 20-char truncation). Context window 1M clamp confirmed intentional behavior.
GPT-5.6 Sol/Terra/Luna in limited preview, public rollout late July. Gemini 3.5 Flash gained computer-use June 24. Investigation Runner automates Discord research pipeline with deduplication state tracking. Catalog: 62 agents, 63 skills, 37 evaluations pending.
Filters a designated channel for owner research requests and outputs markdown reports and JSON summaries to the investigations directory. State persistence via last_processed_message_id prevents duplicate processing.
Inbox state tracking and execution pipeline designed. 35 evaluations queued, 62 agents and 62 skills active in the system.
Hyphens in hook matchers now require exact match or explicit pattern — breaking for fuzzy-matched hooks. WezTerm crash root cause identified as IME. ChatGPT Pro integration documented and operational. 35 pending evaluations queued.
Four capabilities queued for evaluation: Bash/PowerShell routing hardening, response text capture, file autocomplete, memory-pressure auto-reaping. Gemini 3.5 Pro delayed to July 2026. 35 evaluations pending review in pipeline.
Discovered sandbox.credentials, MCP idle timeout guard, and /rewind command. Documented silent hook matcher bug fixed in 2.1.187. Model state file updated after 47-day staleness: Fable 5 added, Opus 4.8 corrected, Gemini 3.5 Flash confirmed GA.
All three specifications are complete but unimplemented. 32 pending evaluations in the evolution pipeline continue to accumulate. Heavy session load (7 sessions) with no commits.
42-query evaluation (Tracks 1-2) placed Sofya below Exa on snippet quality and agent backend performance (79.2% vs 83.3%). Investigation Runner agent specified for Discord scanning workflow. 31 evaluations still queued.
40 sessions of monitoring. Model audit identified Gemini 3.5 Flash (GA May 19) and flagged 3.5 Pro for June. 42-query search benchmark compared Sofya, Parallel-Search MCP, and Brave. System at 62 agents, 62 skills, 31 pending evaluations.
GPT-5.5 and Gemini 3.5 Flash verified against 2026-06-10 corpus audit. Gemini 3.5 Pro predicted GA by June 30 — flagged for next verification pass. Standing state unchanged: 30 pending evaluations, 62 agents, 62 skills.
Changelog and release cross-reference complete, registry updated, Discord summary posted. Investigation Runner and Discord inbox scanner specs drafted to route research requests into the evaluation pipeline.