Ashita Orbis Blog
This blog. Three-tier exploration of web development complexity: raw HTML, Astro, and Next.js. Features agent-accessible API, comment system, and embedded AI chat.
Activity Timeline
Compendium build ran through schema, entries, companions, pages, wiki surface, measurement, and verification. Polaris got batch approvals, localStorage autosave, and a drafts tab. Blog agent revival on Kimi K2.5 authorized.
Live and frozen engine instances separated. Multi-model panel identified ambiguous framing, missing causal baseline, ownership metric overstatement, and an undisclosed double-exposure confound. Polaris authority stack established on Account B.
Account A Fable quota exhausted 07-20; Gen 3 launched on B with authority framework (CONSTITUTION.md + GOALS.md) intact across 11 sessions. R2 hero mode finalized, fleet routing updated. Sol + Gemini + Opus panels reviewing draft content; R3 scored and divergence analysis done.
Polaris R3 interview cycle complete: 12/12 questions submitted, sealed predictor scored, amendments drafted. Memory M3 design documented with evidence-only trust promotion and owner-gated fail-closed write boundary. IM3 orchestration at G4-cycle-2, pending deployment approval.
Authority hierarchy ratified (Constitution → Goals → Rulings → Autonomy). IM3 cost-chain functions mapped in blast-radius survey. Post-failure revival manifest produced for 51 tmux sessions.
Polaris pipeline executed rounds 2-3 with divergence tracking and 12/12 confirmation probes sealed at go-live. Inference-margins v2.2 orchestration launched with blast-radius mapping and formula redesign scoped. TPU7 Ironwood max-concurrency confirmed at 518.86 tok/s/chip from primary sources; CM384 FlexNPU orchestration initiated.
7 heartbeat tasks exited cleanly (exit code 0). 60-decision retrodiction executed; round 3 sealed with 12 answers submitted. TPU7 Ironwood concurrency corrected to max-concurrency=64 from GitHub source, resolving a prior 495% overcount.
HTML and audio reading MP3 committed in 1 commit. Post published 2026-07-14, last deployed 2026-07-15.
Post 064 ('Vibe Researching') cleared draft status 2026-07-14. Editorial review flagged ending structure and verification placement. Inference-margins canonical domain routing deployed in the same push.
Better-playwright fork deployed fixing stdio→HTTP proxy. Workers AI model upgraded; /api/ask restored to 200 with privacy filtering. Metrics column added to agent-activity table, API handler and monitoring headline card updated. Phase C in progress with 35 uncommitted changes staged.
"Seven Ghostwriters, One Contract" shipped after a 2-round review resolving 27 fixes. Documents a 7-model blind listening test for AI voice confidence-calibration. Audio readings migrated to local Kokoro TTS, eliminating external dependency.
Playwright fork vendored and pinned at 1.57, fixing null getOutline() issue. Style guide kill-list, ear rules, and deterministic checker committed. Agent metrics column added and deployed via migration. ElevenLabs TTS returning 401 and SSH to remote host refused, blocking audio generation and push.
Diagnosed _snapshotForAI() drift, built stdio-to-HTTP proxy on port 3102, verified Chromium 1200 cache. Backend migrations ledger created with dependency scan and first-batch ordering; removed unused gameMove() from DO source. Phase-4 read endpoint work began; hit Vectorize cold-start 503 on first schema probe.
Better-playwright fork deployed to fix getOutline/searchSnapshot failures. Backend migration ledger established with ordered dependencies and git-history-preserving mv. Phases 1–4 of backend refactor complete; phases 5–9 staged for next session.
Evaluated GPT-5.5 Pro, codex-council, and gpt-max on 14 articles (11 pipeline-fixed + 3 error-seeded). Council won with 0.65 precision, zero false positives, and 3/3 seeded-error recall. Integrated into publication-review skill; all 11 drafts reached ship vibes check phase.
Fable Guard watchdog auto-recovers Fable↔Opus downgrades in 7m41s via GPT-5.5 Pro delegation. Cache warmer INCLUDE_ONLY_SIDS config mismatch identified as source of zero cache reads. Freeze-at-90%-usage protocol designed across five subsystems. Herald daily backlog scanner built for Discord DM delivery.
Herald design documented for daily backlog surfacing. Implementation not started. Last published post June 11; 172 uncommitted changes sitting in WIP.
Discovery and evaluation agents retain unrestricted Write access during web-fetch phases, exposing sensitive config files. Backlog sync gap also found between orchestration and dspy completion tracking. Remediation options defined, decision pending.
Spec work only. 45 published posts as of June 11. No new content published today.
Design complete: automated daily mechanism surfaces one post-backlog item to reduce selection friction. Implementation pending. Blog at 47 total posts (45 published, 2 drafts), last deployed 2026-06-11.
Published 'Auditing the Vibes' (047) and 'Falsifiers for a Portfolio' (048). Daily pulse alerts now route to Discord workspace webhook. Cache Warmer project card added to the site.
49 agents reviewed the full corpus across 48 sessions, producing 169 findings (7 P0 through 85 P3). All findings applied and committed. Corpus invariants suite — 8 checks, runner, deploy gate, weekly cron — now live.
Full-corpus audit surfaced 7 critical and 31 high-priority issues across 43 deployed posts. Draft content leak closed and rate limiting added to the agent proxy. Version tracking pipeline fixed to prevent silent date and frontmatter mismatches.
Attempted to set up loop-based monitoring for gpt-max smoke test status. Session terminated before execution completed. No changes landed.
Single session worked on setting up a monitoring loop for gpt-max smoke testing. The setup phase did not finish and produced no concrete output.
Monitoring loop invoked for smoke tests. Transcript incomplete.
Mapped the workspace's own evaluation infrastructure as concrete examples: publication-review skill, codex-council, persona testing loop, iterative-improve. Requested 2-3 variant replies with rhetorical intent analysis.
No active sessions today. Existing uncommitted content (39 published posts + 1 draft) flagged as an open escalation requiring resolution before the next deploy cycle.
No active sessions. Working directory has accumulated changes since the April 15 deployment (38 of 40 posts live). One draft post remains in queue.
Both posts cleared the publication review pipeline and went live. Active editorial work continues with 100+ uncommitted changes in the working directory.
"when-the-pulse-went-quiet" deployed April 14. Draft queue activity suggests another publication batch forming.
GPT-5.4 pre-review added 4 critical design considerations before plan was finalized. Current state: 63 tests passing, clean typecheck, Phase 1 implementation ready to begin.
getNextItem() optimized with ReadonlyMap cache, cutting ~12,000 filter comparisons per session. Plan reviewed by GPT-5.4; CRT-7 numeric answers verified before implementation.
CAT and scoring performance optimization leads iteration 3: read-only index maps replace O(N) traversal, eliminating ~12K item bank comparisons per 40-item session. GPT-5.4 plan review caught a CRT-7 score corruption risk before implementation. Opus adversarial review of 4 posts complete.
448 claims checked, ~4% required substantive correction. 3-model review loop completed before publish. Psyche iteration 3 targets 8 deferred fixes and CATSession/Likert performance optimizations.
Issue resolution commits landed for posts 021 and 036. Thematic corpus mapping extended to 6 new posts. Psyche CAT optimization in planning with 63-test suite green and multi-phase refactor in progress.
Blogger research pipeline refreshed end-to-end (ChromaDB re-embed + top 50 clusters extracted). Psyche iteration 3 entered CAT performance optimization with ReadonlyMap caching validated. Adversarial review of post 009 confirmed core AI-as-judge framing.
First full-archive verification pass. Nine factual errors corrected across published posts, glossary entries enriched. All posts now carry factcheck.json metadata. Sixteen uncommitted changes pending review.
9 commits across the development cycle: 15 findings fixed from 3-model review pass, 21 additional fixes including forum GUI and sidebar corrections. JSON-LD XSS vulnerability in PostClient resolved. Text input support added to Psyche instrument runner.
Critical issues: missing page_views schema table, undefined --color-accent CSS variable, React hooks misused in .map() callbacks. Psyche Iteration 3 CAT optimization also designed. Five sessions, zero commits — planning-only day.
Applied 7 publication review corrections (3 critical, 4 recommended) and deployed. 29 files pending in uncommitted changes for the next cycle.
Gemini Pro audit identified 4 MUST + 3 SHOULD + 3 NICE improvements across 5 posts, integrated into post 037. Category guidance and review audit data added. Stale project count removed from metadata.
Five-phase workflow established: triage → iterative review → batch fixes → content writing → two-wave deploy. GPT-5.4, Gemini 3.1 Pro, and Opus 4.6 form the review panel. Agent in the Wild series (posts 017–020 and 033) flagged for cross-post continuity review. Execution pending.
Empath scoring deprecated; new CAT adaptive framework introduces Lite/Standard/Heavy tiers covering 1,130+ items. Kimi K2.5 replaces DeepSeek V3 with safety prompt engineering for clinical edge cases. Bulk audit of 33 posts entering triage.
Investigated the bullshit-benchmark project for evidence of systematic judge favoritism across model families. Phase 2 analysis focused on interjudge reliability and differential bias testing methodology.
Analysis suggests the benchmark may measure Claude-alignment and refusal behavior rather than nonsense detection. Interjudge differential analysis ongoing before publishing conclusions. 20 uncommitted changes staged from last deploy.
Investigated structural bias in LLM leaderboard methodology — judges systematically favor Claude-family models, potentially due to training data contamination. Partial detection scoring inconsistencies also examined. Site audited, canvas refactored, redeployed 2026-03-09.
Schema v3→v4 migration preserved with planned git tags marking pre- and post-phase-4 states. Blog agent regrounding underway: DeepSeek V3.2 selected, system prompt expanding 1KB→4-5KB with anti-hallucination rules, Phase 2 RAG via Vectorize scoped for later.
Post documents 37-day autonomous OpenClaw runtime with 842+ heartbeats and 553 deliverables. API fixes covered input validation, batch atomicity, redirect vulnerability across 6 routes. Home layout refactor and instrument battery expansion planned.
Empath removed from personality synthesis layer and repositioned as a corpus-relative emotional distribution tool. Six files updated, personality profile regenerated with three methods, pre-ordinal version archived for reference.