Claude Evolution System
Self-improving AI development environment. Discovers new capabilities, evaluates them against a scoring framework, and integrates approved tools on cron.
- Autonomous capability discovery pipeline
- 5-criterion evaluation scoring framework
- Auto-integration of approved capabilities
- Multi-model orchestration (Claude, Codex, Gemini)
Activity Timeline
All 5 background tasks exited cleanly. Housekeeping approvals unlocked the integration/reattach-approved-backlog branch. Roster: 27 agents, 69 skills, 50 pending evaluations queued.
Pipeline state file lagged global config on Gemini model version — discovery filed. Night-shift TOCTOU race and four additional defects fixed with probe validation and passing test suites. Discord investigation runner architecture specified.
Model confirmed against Google docs July 21: 1M context, $1.50/$7.50/M. Daily heartbeat capability-discovery job configured. Polaris-C (Fable) launched on Account B with full authority stack inherited.
Investigation runner fetches Discord channel messages, filters topics, dispatches research, writes structured markdown reports. Daily integration executor armed on codex-exec transport. Model reference table validated against current GA dates.
50 backlog items processed, 16 techniques added to library index across four sessions. Eight commits covered 2.1.193/195 registry updates, hook-matcher audit, Capframe, and defense-in-depth notes. Gemini 3.5 Pro detected in freshness sweep and queued for evaluation.
15-version gap (v2.1.200–v2.1.214) closed in one session. 9 discovery files generated for pipeline evaluation. 46 evaluations remain in the pending queue.
State file current as of 2026-07-16. 46 pending evaluations, 25 agents, 64 skills active.
Fable 5 added, Opus 4.8 and GPT-5.6 Sol/Terra/Luna registered. Discovery file created for agent/skill references needing model ID updates. Evaluation backlog and pending discovery sweep keep status degraded.
Design day: 8 sessions, 0 commits. Two Discord-integrated agents specified. Decision-making pattern established (AskUserQuestion for all forks). Evaluation loop stalled at 40 pending items.
Hardened approval gate with unified orchestrator safety model. Added backlog test-and-deploy method with Opus rubric and preregistration tooling. 40 items queued for next evaluation cycle.
Model landscape updated with GPT-5.6 variants announced June 26. Agent specs drafted: Investigation Runner for Discord inbox scanning, workspace orchestrator with 5-tier priority matrix. 43 evaluations pending in queue.
Phase 1 of publish model inversion fully designed but held pending owner approval. Portfolio cleanup shipped in 3 commits: README, private mirror backup, registry path and IP fixes. Model version state file found 6 weeks stale relative to CLAUDE.md reference.
ANTHROPIC_BASE_URL export during RC launch was causing proxy 404s on availability checks — stripped from launch script. Allow-list publishing model ready for Phase 2 review. Multi-account isolation verified via CLAUDE_CONFIG_DIR workaround.
RC was blocked by ANTHROPIC_BASE_URL proxy in launch script — patched. Newspaper crontab failed due to missing logs directory — fixed, job 4 completed. ChatGPT Pro MCP failure modes characterized (timeouts, CDP drops, 20-char truncation). Context window 1M clamp confirmed intentional behavior.
GPT-5.6 Sol/Terra/Luna in limited preview, public rollout late July. Gemini 3.5 Flash gained computer-use June 24. Investigation Runner automates Discord research pipeline with deduplication state tracking. Catalog: 62 agents, 63 skills, 37 evaluations pending.
Filters a designated channel for owner research requests and outputs markdown reports and JSON summaries to the investigations directory. State persistence via last_processed_message_id prevents duplicate processing.
Inbox state tracking and execution pipeline designed. 35 evaluations queued, 62 agents and 62 skills active in the system.
Hyphens in hook matchers now require exact match or explicit pattern — breaking for fuzzy-matched hooks. WezTerm crash root cause identified as IME. ChatGPT Pro integration documented and operational. 35 pending evaluations queued.
Four capabilities queued for evaluation: Bash/PowerShell routing hardening, response text capture, file autocomplete, memory-pressure auto-reaping. Gemini 3.5 Pro delayed to July 2026. 35 evaluations pending review in pipeline.
Discovered sandbox.credentials, MCP idle timeout guard, and /rewind command. Documented silent hook matcher bug fixed in 2.1.187. Model state file updated after 47-day staleness: Fable 5 added, Opus 4.8 corrected, Gemini 3.5 Flash confirmed GA.
All three specifications are complete but unimplemented. 32 pending evaluations in the evolution pipeline continue to accumulate. Heavy session load (7 sessions) with no commits.
42-query evaluation (Tracks 1-2) placed Sofya below Exa on snippet quality and agent backend performance (79.2% vs 83.3%). Investigation Runner agent specified for Discord scanning workflow. 31 evaluations still queued.
40 sessions of monitoring. Model audit identified Gemini 3.5 Flash (GA May 19) and flagged 3.5 Pro for June. 42-query search benchmark compared Sofya, Parallel-Search MCP, and Brave. System at 62 agents, 62 skills, 31 pending evaluations.
GPT-5.5 and Gemini 3.5 Flash verified against 2026-06-10 corpus audit. Gemini 3.5 Pro predicted GA by June 30 — flagged for next verification pass. Standing state unchanged: 30 pending evaluations, 62 agents, 62 skills.
Changelog and release cross-reference complete, registry updated, Discord summary posted. Investigation Runner and Discord inbox scanner specs drafted to route research requests into the evaluation pipeline.
/config key=value syntax is the highest-priority find (score 82), enabling settings injection in headless scripts. Critical bugs: prompt caching broken on custom base URLs (silent cost inflation), Write/Edit producing 0-byte files on network drives (data loss risk).
State file logged 2.1.177 while CLI was at 2.1.176 — caused by a prior RSS run. Both versions already documented; 2.1.177 is compliance-only. New specs written for Discord #general scanning and automated topic research (Investigation Runner). Workspace orchestrator priority matrix formalized.
US export control directive (June 12) triggered the patch. Investigation report filed, capability registry updated with model_suspension flag. Investigation Runner agent spec drafted for automated future monitoring.
Language setting for session titles and footerLinksRegexes discovered as novel additions. Pipeline queued 2 evaluation specs. Gemini 3.5 Pro identified as new frontier reasoning model; state update initiated but incomplete.
Six novel capabilities identified including enforceAvailableModels governance feature. Model table updated with latest external models. Four Discord workflow specs drafted for autonomous research pipeline.
30 changes classified across v2.1.170→2.1.172; 3 discovery files created. Workspace audit redesign proposes Discord reaction-based approval flow for ranked action queues. 25 evaluations still pending.
claude-fable-5 registered with always-on thinking and $10/$50 pricing — flagged as cost concern for lightweight tasks. Gemini 3.5 Flash (~70% cheaper than 3.1 Pro, GA June 8) queued for visual analysis evaluation. VS Code transcript bug resolved.
30 changes classified across the v2.1.166→v2.1.169 window, including 3 novel features: --safe-mode flag, /cd command, and disableBundledSkills setting. Two new Gemini 3.5 models added to the discovery registry and contemporary models state refreshed. 22 evaluations pending human review.
version-tracker.sh bug caused apparent v2.1.166↔2.1.168 churn; both patches confirmed as stability-only. Two new Gemini models cataloged from model verification check. Investigation Runner agent spec drafted but not deployed; 20 evaluations remain in backlog.
GPT-5.6, Gemini 3.5, Gemini 3.5 Flash, and Gemini Omni catalogued with discovery files. v2.1.168 registry entry updated as bug-fix release. Seven sessions spent specifying an automated Discord research intake agent; specification complete, no implementation yet. 19 evaluations pending.
Novel finds: deny-rule glob patterns, MAX_THINKING_TOKENS=0 for disabling thinking, cross-session messaging authority isolation. Gemini 3.5 Flash and Gemini Omni Flash identified as visual-inspector upgrade candidates. Registry, pipeline state, and Discord updated; state file verification date blocked by permissions.
Delta fan-out audit approach (10-15 agents, emoji-reaction approval) approved and deployed. Claude Code v2.1.163-165 documented. Agent gateway auth failures under investigation post-GPT-5.5 cutover. 15 capability evaluations queued; investigation runner agent specs drafted.
Three sessions investigated v2.1.159–v2.1.162 changelog. Breaking keyword rename flagged HIGH priority for CLAUDE.md update. Gemini 3.5 Pro launching June 2026 caught as stale state entry. Registry updated, 3 discovery files created, Discord notified.
Version check found no user-facing changes in v2.1.159 — no agent or skill registry updates needed. Contemporary models audit surfaced Gemini 3.5 Flash as a new untracked model; discovery file created and reference docs updated. Analysis posted to Discord.
Model verification found two Google releases (May 28-29) newer than the tracked config entry. State file update blocked by permissions; discovery routed to pipeline. Investigation Runner agent architecture drafted for Discord monitoring integration.
Executed INVESTIGATE-UPDATE.md playbook across 16 Claude Code versions. v2.1.154 flagged as major release introducing Opus 4.8 and Dynamic Workflows. Registry update pending write approval; 10 evaluations now in backlog.
Heartbeat surfaced stale tracking: v2.1.153 shipped day-of with a /doctor diagnostic command targeting stale trigger loops. Model reference updated for Gemini 3.5 Flash (Google I/O GA). 9 capability evaluations pending in evolution queue.
Version delta v2.1.140→v2.1.152 classified as substantial and added to capability registry. Gemini 3.5 Flash (Google I/O, May 19–21) queued for evaluation. Specification designed for a Discord #general monitoring and investigation-report automation pipeline.
CLI 2.1.140 vs state file 2.1.150 mismatch causes version-check false positives every run. Root cause documented, fix pending. Discord inbox router and investigation-runner agent specifications written for next implementation cycle.
state/versions.json had current/previous swapped, causing false-positive version detection. Investigation report written, Discord notified. Contemporary models table verified. 6 pending evaluations and the state file repair remain open.
Auto-updater stalled ~May 12. v2.1.126-128 features (OTEL, workspace server) present in binary but missing from npm registry — flagged as unreleased artifacts. Feature registry and model docs updated. Discord Investigation Runner specs drafted.
Four Claude Code releases documented with feature extraction and bug triage. Google I/O 2026 model check added Gemini 3.5 Flash and Gemini Omni Flash to the evaluation pipeline. Registry updated to 62 agents / 62 skills.
Novel additions include claude agents --json interface, /plugin preview, mouse support, and OTEL agent_id for distributed tracing. A /simplify vs code-review:code-review naming conflict identified for follow-up. Model verification pass confirmed no new GPT-5.5 or Gemini 3.1 Pro releases.
Workflow designed to monitor Discord #general, run research via WebFetch/Grep, produce structured markdown reports, and post Discord embed summaries. Design phase only — not yet executing.
Downgrade caused by inverted NEW/OLD params in version-tracker.sh. Downgrade report committed to state, Discord alert posted at orange severity. Model registry verified current; semver downgrade guard flagged as recurring infrastructure debt for the second time.