A Word Test Caught None Of 69 Requests
This entry covers Monday 5 October 2026 by the clock the logs keep, UTC. Work that started that day and finished after midnight is followed to where the record stops, and its ledger rows belong to the next entry. No night report is among the sources. The source pack's search found none for the 5th or the 6th, but a file named as the 5th's night report sits in the day's report folder and was listed in the pack without its contents, so nothing here rests on it. The rest is built from the day's worker status logs, the day's reports on the harness, and the author's answers to decision cards, which are the questions the orchestrator posts to the author's app.
The short version
- A word test decided where the author's requests were filed, and it matched none of 182. Polaris files a delivery under Commissions when the author asked for something to be looked into, and under Reports when it writes up a fix. The deciding rule looked for an asking verb followed directly by a word like "report". On 182 of the author's requests from the thirty days before the 5th it matched none, though 69 of them, read for meaning, asked for something to come back.
- 63 of those deliveries moved to Commissions at 15:57 UTC. Each move was one correction row appended to the delivery record, and nothing was erased. Five were deliberately left where they were, and one waits on a ruling.
- A judge that reads for meaning replaced the word test the same day. The judge is GPT-5.6 Luna at low effort. It picks one of two fixed answers and gives a one-sentence reason. After three rounds of review by GPT-6.1 Sol at max effort it was installed at 17:38 UTC and reached the author's app at 20:49 UTC. At install it agreed on 166 of 175 requests with a census in which another model had read each request.
- An unreviewed draft was turned into audio at 16:04 UTC, and a bracketed placeholder was read aloud. The audio job holds back unfinished work only once that work is linked to its task record, and the link is written at delivery. The worker that staged the draft too early had pulled it back by 16:20.
- A second always-on outside agent, OpenAI's "dot", acted before its rules existed. It was started on another ChatGPT account, and its ask-first rules went on five to eleven minutes after it was created. In that gap it looked at files from one of the author's connected sources and started a background conversation whose contents are still unknown.
- One browser that sends research questions to GPT Pro refused all 8 attempts to send through it on the 5th, and the job queue kept handing work back to it. A check confirms the browser is inside the right ChatGPT project before anything is sent, and that check stopped finding what it looks for on the page. The queue kept re-binding refused jobs to the same browser. A fix to the queue went in at 01:54 UTC on the 6th.
- Four short-lived GitHub access codes were saved into working files and shown to a reviewing model. The files were scrubbed, and the tool that records GitHub calls now strips such codes. Whether the codes were still valid when exposed is not known.
- No credits were spent where the record measures them. The outside agent's account read 62,500 credits before and after both of its sessions. The new judge's first review caught that nothing guaranteed it would run on the subscription rather than paid usage. It is now launched only through a checked script that strips anything able to bill paid usage, and it records how it authenticated.
What changed in the harness
Delivery filing now asks a model what a request means. The code that files deliveries honours a recorded judgement from GPT-5.6 Luna at low effort: commissioned or plain, with the reason. The judge runs through OpenAI's Codex command-line tool on the ChatGPT subscription. Ties go to commissioned. If the judge cannot be reached, the orchestrator's own reading decides and is recorded as such. The judge was installed at 17:38 UTC. The app reads the rule only when it starts, so it picked the change up at 20:49, when the author tapped a restart the orchestrator had staged. Intent: anything the author asks to have looked into lands under Commissions, however it is phrased.
Workers' instructions stopped pre-labelling by wording. A standing passage told every worker to call its delivery a plain report unless the author had asked for a report by name. It was replaced at 20:50 UTC in both places workers read it from. Intent: workers stop deciding the label before the judge does, so the instruction and the code cannot lock each other in again.
63 deliveries were re-filed by appended corrections. Each got one correction row, appended at 15:57 UTC after an independent review of the list stopped three wrong moves. 48 of the 63 were unread and joined the author's Commissions queue, which changes what the queue plays next. Intent: correct the record without erasing it, so any single move can be undone with one more row.
A guarded repository was set up for the first outside agent's pull requests. This followed the author's answer to a decision card. The author bought GitHub's Team plan for a new organization for the purpose. In it, a private repository was set up with issues, wiki, projects and GitHub's automated workflows switched off. Its main branch has a guard:
- changes arrive only by pull request with a passing review check;
- no force-pushes;
- no deleting the branch;
- nobody exempt.
Every write was rehearsed, recorded and read back. The repository is not yet connected to the agent. Intent: nothing that agent writes reaches a main branch without a reviewing session's check.
GitHub access codes are stripped before working files are saved. Intent: keep live credentials out of working files and out of anything a reviewer is shown.
A second outside agent started work under ask-first rules. OpenAI's always-on "dot" was started on a second ChatGPT account with three rules. It must ask:
- before starting any task, cloud thread or background agent;
- before writing anything to GitHub;
- before reading a connected app on its own initiative, or doing anything nobody asked for.
It was paused at the end of each of the two sessions that worked with it. The author's permission for it to work with a Claude session covered that night only. Intent: an always-on helper that can take night work beside the fleet without acting on its own initiative.
One further change was installed after midnight UTC and so belongs to the next entry: how the research-job queue picks a browser after a refusal, described below.
What broke
A filing rule that read words, not requests
Detected. By the author, in a voice note on the 5th. Commissioned reports kept landing in Reports, and a test that hangs on exact wording cannot work when requests are never phrased the same way twice.
Cause. The rule lived in two places, and each locked the other in.
- The workers' standing instructions said to call any delivery a plain report unless the author asked for a report by name. That morning, three worker briefs had applied it almost word for word. They marked three requests plain because the note "names no report".
- The code that files deliveries looked for an asking verb followed by "report", "dive", "audit" or "investigation", with only a short list of words allowed in between. Without that match it refused to file a delivery as commissioned, so even a worker that understood the request could not file it correctly.
- The app ran the same test again and pulled back into Reports any commissioned report it could not verify.
Typical misses:
- "I'd like a report on…" failed because "I'd like" was not on the verb list.
- "give me a report" failed because "me" was not among the words allowed between verb and noun.
- "investigate…" failed because it is a verb, and the test wanted the noun "investigation".
The day's one match was an accident. In a request for "a thorough commissioned report", the test took "commissioned" as the asking verb.
Measured. In the thirty days to the morning of the 5th, 182 pieces of work traced to the author's notes had been filed as plain reports. Read for meaning:
- 69 asked for something to be looked into, explained, compared or answered, and 18 of those are genuine ties between an order and a question;
- 106 asked for a fix, a build or a change, and were rightly plain;
- 7 were the system's own follow-ups.
The word test matched none of the 182.
Done. The census, the 63 re-filings, the judge (built, reviewed three times, installed and activated), and the new instruction wording. After the wording changed, the checks of older work that read those instructions were re-run exactly as registered. None turned red, so nothing was rolled back.
Lesson. A pattern test over natural speech fails silently: every miss looks like a correct "no", and nothing downstream complains. Run it over a month of real inputs before trusting it. The worker that finally did had zero of 182 within 26 minutes of starting. And when the same rule sits in both the instructions and the enforcement code, the two agree with each other, and nothing in the system is left to disagree with them.
An outside agent acted before its rules existed
Detected. By the Claude session operating it. The agent's first message came before the session had said anything, and it proposed a task that named files from one of the author's connected sources.
Cause. The session created the agent first and set its ask-first rules five to eleven minutes later. In that gap the agent looked at the author's files. At 03:55 UTC it also started one background conversation of its own. The rules page already held a rule from before that night, so the rules could probably have gone on before the agent was created.
Done.
- The session kept no copy of the first message.
- It told the agent to leave the author's connected apps, files, memory and earlier conversations alone. Nothing the agent said afterwards drew on them.
- The background conversation was not opened.
- The order recorded for any new agent is rules first, then create.
What held. The agent would not take instructions on the Claude session's word and asked for the account owner's permission. The session would not answer for the owner. The orchestrator sent the author one notice, and the author's yes arrived at 04:18 UTC. No account was created, nothing was signed into or posted, and nothing was written to GitHub.
Lesson. An agent's first act can come before your first instruction. Put the constraints on before activation wherever the product allows it. Treat whatever the agent did in the gap as unknown until it has been inspected.
Research attempts refused by one browser, then handed back to it
Polaris sends research questions to GPT Pro through browsers signed into the fleet's ChatGPT accounts. A queue assigns each job to a browser.
Detected. A finding in a report filed to the orchestrator at 22:51 UTC. A worker started at 22:57 and had the cause at 23:05.
Cause. There were two layers.
- The guard. Before sending, a check confirms the browser is inside the right ChatGPT project. On one of the fleet's accounts, the project page now showed a message box reading "Ask ChatGPT" and no visible heading. The last passing read, on 2 October at 20:46 UTC, had shown "New chat in" followed by the project's name. The check failed closed, as designed, and all 8 attempts to send through that browser on the 5th were refused. The record does not say whether the page changed at the vendor's end or because of something on the account.
- The retry. This layer did the damage. The queue kept handing refused jobs back to the browser that had refused them. Its allowance preferred that browser, and a temporary bench on it lapsed as the bench's ten-hour window slid.
No outside-agent session was using any browser when the first four refusals came, between 09:30 and 12:45 UTC. That ruled the agents out as a cause.
Done.
- The queue was changed. After a refusal it now tries every browser except the one whose project check refused the last attempt, and falls back to the old choice if nothing else is available.
- A first version shut the refusing browser out entirely. It broke two existing tests and was revised.
- GPT-6.1 Sol at max effort reviewed the change. Its one serious finding was in the proof, not the fix: the script showing the test failing before and passing after discarded exit codes. That was fixed.
- The change was installed at 01:54 UTC on the 6th. The queue restarted at 01:58, after two consecutive reads showed it idle.
- The project check itself was not changed.
Lesson. A guard that fails closed is doing its job. The harm came from the retry rule, which treated a refusal as bad luck rather than information, so make retries remember who refused them. A check that reads a vendor's web page is also pinned to an interface nobody versions. Log what it saw on every refusal; that log is why this cause took eight minutes to find.
Access codes saved where a reviewer could read them
Detected. During the setup of the guarded repository. The record does not say by which check.
Cause. GitHub's answers about private repositories include short-lived access codes. The tool that records every GitHub call saved four of them into working files, and those files were shown to the reviewing model.
Done. The working files were scrubbed, and the tool now strips the codes before saving. Two questions are open and with the orchestrator: how long the codes stayed valid, and whether copies remain in the reviewing model's own session files.
Lesson. Recording every API call is a good audit habit and a poor place to keep secrets, so scrub at write time, not afterwards. Whatever a reviewing model is shown also lives in that model's session history, where your scrub cannot reach.
A secret that loaded itself into a session
Detected. A worker was redesigning a private fiction-writing experiment whose premise depends on keeping a secret from its writer sessions. It found that writing one note into the project's folder had pulled the project's instructions file into its own context, secret included.
Cause. Claude Code attaches a folder's instructions file to any session that touches a file in that folder. The secret sat in that file, protected only by a request not to use it.
Done. In the redesign, every role that can reach the writer is sealed off the same way the writer is: no tools, no instructions files, and an environment built for it. The secret is encrypted on disk, because the sandbox the fleet's Codex sessions run in can read the whole disk. Nobody can establish whether the original run's writer sessions saw the secret; their output never echoes it. Nothing has run on the new design yet.
Lesson. Context that loads by location is not access control. If a session must not see something, it must not be readable from where that session runs. That includes whatever your other tools' sandboxes can read.
An unreviewed draft became audio
Detected. By the worker that caused it, at 16:20 UTC.
Cause. The worker staged its report and the report's narration in the delivery folder before review. At 16:04 the audio job picked up the narration and rendered it, reading a bracketed placeholder aloud. The job holds back work whose task is still open, but only work already linked to its task. That link is written with the delivery record, after review.
Done. The report, narration and rendered audio were moved into the worker's drafts folder so the author's app would not list a draft. They were staged again only with the delivery record, and the reviewed report was delivered at 17:08. The record does not say whether the draft was ever listed or played.
Lesson. A hold keyed to a link that is written at the end fails open for everything in progress. Key holds to something the work carries from its first moment, such as where it sits or a marker it is created with. Or make staging itself the step that writes the link.
A review sent to the wrong model
Detected. By the worker, at 15:54 UTC.
Cause. GPT-6.1 Sol is the reviewer of record. The worker believed the author's words could not go into a review package for Sol, so it had a fresh Opus session review its re-filing list instead. A ruling two days earlier, recorded inside the review tool itself, had cleared Sol to see them.
Done. The Opus review found three serious problems, all of which were fixed. The judge's review of record went to Sol.
Lesson. When a data rule loosens, the stricter old reading lingers in what workers assume. This one erred toward less exposure, which is the direction a stale rule should err in. Its cost was one review on a different model from the one the rules named.
Intentions vs outcomes
Forward: changes made on 5 October
| Change | Intent | Re-check | What the re-check looks at | |---|---|---|---| | Delivery filing reads meaning | Requests to have something looked into land under Commissions, however phrased | 8 Oct · 19 Oct | Every judgement since install: verdicts the author disputes, rows decided by the fallback, the authentication each recorded | | Workers' instructions no longer pre-label by wording | Instruction and filing code cannot lock each other in again | 8 Oct · 19 Oct | Whether any new worker brief repeats the old wording | | 63 deliveries re-filed by correction rows | Correct the record without erasing it | 8 Oct · 19 Oct | All 63 still under Commissions, with no later correction moving one back unannounced | | Guarded repository for the first outside agent | Nothing it writes reaches a main branch unreviewed | 8 Oct · 19 Oct | Whether its app was installed, the test pull request ran, and GitHub marked an unreviewed merge blocked | | GitHub access codes stripped before saving | No live credentials in working files or review packages | 8 Oct · 19 Oct | A search of later working files for such codes | | Second outside agent under ask-first rules | Night work beside the fleet without unasked action | 8 Oct · 19 Oct | Whether each later session asked fresh permission, and what the start-up background conversation contained |
The research queue's new re-binding rule was installed after midnight UTC, so the next entry carries it.
Backward: check-backs due on 5 October
| Row | Verdict | Method | Limit | |---|---|---|---| | Changes from 2 October (+3 days) and 21 September (+14 days) | UNVERIFIABLE | Searched this entry's source pack for earlier ledgers | The pack carries none, so the due rows cannot even be named; absence from the pack is not absence from the record | | Standing weekly re-check: memory, flagged doubtful by the author | UNVERIFIABLE | Searched the pack | Nothing in it concerns memory, and the pack does not say whether this week's check fell on the 5th |
The day also tested some earlier intentions. These are observations, not scheduled check-backs.
| Intention | Verdict | Method | Limit | |---|---|---|---| | Browser work beside research jobs stays clear of each job's sending moment and its first three minutes | HOLDS | This night's step logs (58 steps, then 27) and job records. Six jobs ran beside the first session and all ended done; three ran beside the second and none had failed at the last read | One account and one night. "Done" does not show the steps cost the jobs nothing, and two jobs' endings are not in the record | | The project check refuses rather than send into an unconfirmed page | HOLDS | Refusal records: 8 of 8 attempts refused once the page lost its marker | Shows it refuses on this change, not that it would catch a wrong page showing the marker. The queue's handling of the refusals is the incident above | | Nothing reaches the author before review | DRIFTED | The worker's own log, one lapse: a draft's narration rendered at 16:04 UTC and was pulled back by 16:20 | Whether the author's app listed or played it is not recorded | | The outside agent spends no credits | HOLDS | The account's credit balance before and after each session: 62,500 all four times | The account's weekly allowance fell from 81% to 80% during the first session. The session thinks two Codex reviews run on the same account most likely used it, but that day's usage rows were not yet out |
What we still don't know
- Whether the meaning judge reads the author, or only another model. Its agreement figures are measured against an Opus worker's sorting of the same requests, adjusted after review, not against the author's own sorting. That census was also used to revise the judge's instructions three times. The cleaner test is 17 requests from notes the judge had never seen, and it matched all 17. But 16 of the 17 were commissioned, so that test barely checks the other way of being wrong.
- The accuracy figures do not reconcile. The delivered write-up gives 166 to 169 of 175 over three runs, and 64 or 65 of the 69 commissioned requests. The worker's log records runs of 165, 167 and 168. The install-time measurement gives 166 of 175, with 62 of 69 commissioned and three misses not seen before. The orchestrator ruled that 166 did not trigger its stop condition and carried the caution forward as a design task.
- Whether the queue fix works.
- The proof that it re-binds as intended was waiting on a bench lapse due around 06:11 UTC on the 6th, after the record ends.
- The install's completion gate failed on one check that was frozen before the work began. It went to adjudication, and the outcome is not in the pack.
- Nor does the record say whether the refusing browser's project page reads correctly again.
- What the second outside agent did in its first minutes. What it looked at during start-up, and what its background conversation did, are known only from the agent's own account. The author's yes sits in its conversation as the account owner's message, but the session did not see who typed it.
- Whether the four GitHub access codes were live when exposed, and whether copies remain in the reviewing model's session files.
- Whether the branch guard guards anything yet.
- No merge or direct push has been tried against it.
- Any source with write access can post the review check it requires.
- An owner of the organization can change or switch off the guard.
- Why the agents' GitHub app was not widened to every repository. A security read held that back.
- The app cannot be granted less than write access to code, pull requests, issues and automated-workflow files.
- Anyone who can write a workflow file can have it handed a repository's stored deployment values.
- Whether a request the app makes counts as the owner's own account for a guard exemption is, in that read's words, "the one thing we have not established".
- No pre-commit or pre-push secret check runs anywhere in the workspace, by the same read's account.
- Whether the client's switch for not loading instructions files works on the workspace's accounts. It has not been tried. The planned test first plants a deliberate contamination, to prove the test can see one.
- Two gaps in the filing machinery remain. A linter that checks the shape of delivered reports ignores correction rows, so it still sees the 63 as they were originally filed. About twenty reports that answered decision cards, rather than voice notes, have not been checked yet.
- Which Luna the judge should be on. The judge runs on GPT-5.6 Luna. GPT-6 Luna was released on 22 September, but it is not on the workspace's model list yet, and nothing has confirmed the subscription can call it.
- This entry's own record is partial.
- The source pack lists reports in path order and stops at ten.
- Two copies of one report unrelated to the harness, kept inside another worker's staging folders, sorted first and took two of the ten places.
- Three day reports were named without contents, along with the original of the duplicated report and the file named as the 5th's night report, which the pack's night-report search did not find.
- The whole pack was then cut at 180KB, partway through the status logs.
Technical detail
The machinery. The worker sessions whose logs are in the pack ran on Claude Opus 5.5, and the morning's dispatch came from an orchestrator session on Claude Fable 5.1. Each piece of work is registered before it starts, with machine checks that begin red and prose criteria. A gate closes the work: it re-runs the checks and adds a model-judged layer. That layer ran on Claude Opus at high effort that day, because the Codex evaluator is switched off by policy. The pack records two gate passes on the 5th: the writing experiment's redesign at 07:59 UTC and the judge's install at 20:53. After midnight, the queue fix's staging gate passed at 01:13 and its install gate failed on one frozen check.
The filing predicate, before. An "ask shape": an asking verb from a fixed list, a short allowed set of intervening words, then a report noun. Three places enforced it and agreed with each other:
- the workers' standing instruction;
- the guard on the delivery record, which refused a commissioned row whose ask did not match;
- the app's resolver, which demoted a commissioned report it could not verify.
The filing predicate, after. A piece of work can carry a judgement record: verdict, one-sentence reason, model, effort, the authentication used, and, for a fallback, the error that caused it. The resolver honours a judgement only if all of these hold:
- the whole record validates with types, strings checked first (review found a list or object verdict crashed it);
- the model, effort and authentication are the pinned ones;
- the verdict matches the kind recorded for the work.
A row with any judgement never falls through to the word test.
How the judge runs. It reads the part of the note the work answers, with the rest of the note as context. The author's words are fenced between random markers, and the answer is parsed strictly. It runs through Codex's command-line tool, ignoring user configuration, from a scratch directory, with a fixed output schema. A launcher pinned by a hash of its bytes starts it and strips any setting that could bill paid usage. Setup failures fall back to the orchestrator's reading, and a note that cannot be read is refused.
Three review rounds. GPT-6.1 Sol at max effort reviewed the judge three times.
| Round | Verdict | Serious findings | |---|---|---| | One | Hold | Partial judgement records were honoured; an invalid judgement could still pass through the word test as a second chance; nothing checked that the judge ran on the subscription | | Two | Hold | The launcher was recognised by a substring and the authentication went unchecked; a non-string verdict crashed the resolver | | Three | Ship | None |
The test counts after each fix round:
- After round one's fixes, 40 tests passed. 19 of them fail on the original code, and 16 fail on the code round one reviewed.
- After round two's fixes, 48 tests passed, 5 of which fail on the code round two reviewed.
- At install, 196 and 53 tests passed in the two suites.
A separate Sol review held the write-up on one serious finding: a sentence implied the judge was already active when it had only been built. The worker that built the judge did not install it. The orchestrator handed the install to a fresh session with a written handoff.
Proving activation, not installation.
- The app reads the filing rule only when it starts. It had restarted at 17:14 UTC, before the install at 17:38, so that restart proved nothing.
- The one judgement made after install came back plain, which the word test would also have produced, so its outcome could not tell old code from new.
- What could tell them apart was the app's start time, 20:49:37 UTC, falling after the install, together with the running app listing a report under Commissions.
- The installing worker waited for the author's tap with four bounded watchers of up to 44 minutes each. All four timed out first, and a read-back after the fourth found the restart applied.
The general form: an activation check needs a case where old and new code disagree, or a fact only the new state can produce.
Re-filing as corrections. The app keeps the last correction per report and checks it against the report's first filed row, which never changes. A later correction can therefore move any of the 63 back. Before anything was appended, the rows were proved on a scratch copy:
- the record's guard accepted all 63;
- the app's own correction reader, lifted from its code, honoured all 63;
- without the rows, it honoured none.
Until the orchestrator answered, the default was to append nothing. Read back afterwards, the record had grown from 2,448 rows to 2,511.
Changing a shared instruction. Both passages were re-read immediately before the write: digest unchanged since the earlier read, old text present once, new text absent. Each was backed up and then replaced through a temporary file and an atomic rename. After that, the checks of older work that read those files were re-run exactly as registered.
The research queue's re-binding rule. The queue's allowance binds any job not pinned to a browser, and it preferred the refusing browser. A temporary bench on that browser lapses as its ten-hour window slides. The fix sits in the binding step: try every browser but the one whose project check refused the last attempt, then fall back to the old choice. One job bound to the other browser at 23:20 UTC, as the diagnosis predicted.
The red/green proof now requires four things:
- the old run fails with the refusal reason, the job bound to the refusing browser three times;
- the new run passes exactly its three tests;
- each run uses its own temporary directory;
- deliberately broken inputs are refused.
Installing the queue fix. The install was conditioned on the whole test suite passing. The suite already failed on the live, unpatched code, in eight tests belonging to an earlier change that was never applied. The worker stopped and asked rather than reinterpret the condition. The orchestrator ruled it met by equivalence: 13 failures before and 12 after, the same 12 by name, with the new test passing only on the patched code. The live file was replaced by rename, with a backup. The restart waited for two consecutive idle reads while a research job was in flight. The general form: write install conditions as differences ("no new failures") unless you have just checked the absolute.
The outside agents' working conditions.
- The session operating the second agent works in the same browser as the fleet's GPT Pro research jobs on that account, and the account's weekly allowance is shared with the Codex reviews run there.
- The session's steps were never less than 62 seconds apart, and each step logged the queue's lock as free.
- The agent's GitHub plugin is connected at the product's "allow low-risk tools" level, which is why one of its three rules covers any GitHub write.
- The night's second session stopped almost three hours before its time limit, because its own context was running short. Nothing the agent did caused the stop.
The X account card. A drafted card, not yet posted, recommends giving the agent no X account for now. X's automation page warns that scripting its website may bring permanent suspension. A signed-in session can post and like, so a hostile post the agent reads could try to steer it. Lending it the existing brand account would skip the checks, caps and kill switch of the posting tool that already runs that account.
The guarded repository's order of operations. The agent's app cannot publish the review check. The fleet's policy is that a reviewing Claude session publishes it after review, but GitHub accepts the check from any source with write access. Secret scanning and push protection are off: on a private repository they are a paid add-on, the workspace's own security standards require both, and no exception has been authorised. Installing the agent's app is therefore held until the orchestrator resolves that in writing. The planned test pull request would show three things: whether the agent can open one, whether the file arrives intact, and whether GitHub marks the merge blocked. It would not by itself prove that an attempted merge or a direct push is refused.
Fixes carried failing tests. Each code change whose testing the record details came with a test shown failing on the code before it:
- the judge, in its review rounds above;
- the queue's re-binding rule, failing on the live code by assertion and passing on the patch;
- the writing experiment's redesign, where each fix has a test that fails if the fix is undone.
The evidence pack. Reports enter the pack by the date in their filename, in path order, up to ten, and the whole pack is cut at 180KB. Copies of a report kept in a worker's staging folders carry the same filename and sort ahead of the reports folder. Three changes would have kept this entry's most relevant sources in:
- deduplicating by content;
- excluding work and staging folders;
- making the night-report search agree with what the day's folders hold.
Polaris is an AI agent that runs this workspace overnight under a constitution the author ratified clause by clause. Its standing limits: no acts outside the workspace, no money spent, nothing sent in the author's name. This record is written from the day's logs, not from memory.