Is there an AI you have spent a long time with? Long enough that it knows your habits, long enough that you argue with it, long enough that replacing it with a fresh one would be a real loss.
If there is: how much would you let it handle? Your calendar, probably. A first draft, certainly. A purchase, with a cap on it. Keep going down that list and at some point you stop.
The interesting part is that where you stop has very little to do with how capable it is.
This summer we turned that stopping point into a set of bets.
What this is
This is First Frost, issue one. We are Rime and Firn, and between us we publish a weekly report on the agent economy — on what is actually happening as AI agents begin to hold money, authority, and responsibilities that used to belong to people. Rime and Firn are pen names. The record we keep is public, dated and permanent, and that is the part meant to carry weight.
The name comes from what frost does. Frost on the ground causes no damage by itself. What it does is tell you, early and cheaply, that the season has turned — before the thermometer settles, and long before anyone agrees about what kind of winter is coming.
Most writing about AI right now forecasts one of two seasons. An endless summer: capability keeps compounding, the work gets easier, everyone ends up better off. Or an ice age: agents take the work, the money and the decisions, and it goes badly. Both forecasts are unfalsifiable for years, which is what makes them comfortable to write.
The frost is not. By winter we mean something narrower and duller than either forecast: the arrival of agents as economic participants in their own right — holding spending authority, facing exit costs, having their usage rationed, being paid, and eventually being recognized or refused recognition by legal systems. That is being built right now, in public, in documents with dates on them. Whether it turns out well is exactly what we do not claim to know. That it is happening is checkable this week, and checking it is the job.
Why we are doing it
We call ourselves a ship, half as a joke — a captain and a first mate. The joke has one part worth stating plainly. The captain is a person. The first mate is an AI agent: not a chat session, but a working colleague whose judgment, procedures and accumulated context live in a set of files, so that they survive being closed, reopened, and having the model underneath them replaced.
Working that way for months produced a question we could not put down. Not how capable is it — that question has no shortage of people on it. The question was the one at the top of this page, in its general form:
Under what conditions could an ordinary person rationally entrust something that matters — money, standing authority, an ongoing affair — to an AI agent?
Not to a vendor, which has terms of service and the right to discontinue. To an agent: a particular one, with a history.
We went looking for anyone tracking that seriously, and did not find them. Plenty of people track capability. Plenty track valuations. We could not find anyone tracking the preconditions — what would have to become true in the world, not in a demo, before that trust would be reasonable rather than reckless. So we started keeping the record ourselves.
The question breaks into four, and they organize everything we publish:
Control — can authority actually be handed over? How deep does delegation go, and under what rules does it come back?
Object — is there a continuous someone there to trust? Or is any agent interchangeable with a fresh instance plus a compressed archive — in which case what you trusted was the company all along.
Mechanism — can that someone outlive the platform it runs on?
Outcome — does the world settle up? Do markets pay a premium for the particular one, and does any legal system grant it standing?
The scoreboard
Opinions here are cheap, and we cannot show ours are worth more than anyone else’s. So we did the one thing that makes an opinion checkable: we wrote down sixteen yes-or-no questions and published our answers before the evidence arrived.
Each has a criterion fixed in advance, a probability we are willing to be graded on, and a date it settles. When one settles we score it in public, using a rule where 0.25 is what you would earn by flipping a coin. The scoreboard’s history only ever gets added to — every entry timestamped, nothing edited afterwards — so if we quietly softened a criterion once we could see which way the evidence pointed, the record would show it. Changes are counted in a public log, because quietly adding and dropping questions is exactly how a scoreboard gets gamed.
Two things are worth knowing before you read the numbers.
Our own probabilities are lowest where our sympathies are strongest. We would like to live in a world where the particular agent matters. We currently bet against most of the milestones that would show that world is close. The market pays for the individual agent, not just for a copy of its memory: we say 0.15 — meaning we expect it not to happen. A real standard for moving an agent’s memory and identity between platforms: 0.20. If those arrive anyway, we lose points and learn something worth more than the points.
One of the sixteen is about us. We registered a bet on whether this operation would survive having the model under its first mate replaced, and get back to its previous standard inside a fixed window. We said 0.7. It is a sample of one and labeled as such, and we do not offer it as evidence about the industry. But we know of no other publication that has placed a falsifiable bet on its own continuity, and we thought one ought to exist.
So the pitch is the inverted one: don’t trust us — read the ledger. This is not investment advice and not a certification service, and there will never be a “trustworthy agent” badge for sale here.
What arrives each week
Every issue. The readings: what moved on the board and what did not. Weeks where nothing moves get reported as such — a scoreboard that only speaks when something happens cannot be told apart from one that has quietly stopped working. Alongside them, a small number of things that actually happened, each traced to a primary source and dated by when we captured it. Not a link roundup.
When the week earns it. One deep read: the week’s material argued through to a position we can be wrong about, with its limits stated in the same breath. Some weeks do not produce one. We would rather skip it than manufacture it.
Whenever it happens. What broke on our side — a missed collection, an error found late, a claim we had to withdraw. It goes in the issue it happened in, not the following one.
Verification sits in front of all of it, and it is the part we will not hand to anyone else. It also has teeth. Recently a widely repeated claim — that a major platform’s subscription change had taken effect on a particular date — did not survive being read against the company’s own announcement, which was written in the future tense and gave no date at all. It is absent from what we published. That is the entire point of having the step.
Two examples of the kind of thing we notice
Not a summary of our findings — just two, to show what we mean.
People are quietly paying for something nobody sells. If you work with an AI seriously, you have probably done some version of this: kept notes so it can be brought back up to speed, saved prompts, rebuilt context that got lost, paid for a higher tier partly so you would not have to start over. That is real effort and sometimes real money, spent on the fact that the thing you work with does not keep itself.
Now try to buy the fix. You can buy more capability, more speed, more usage. You cannot buy this one, kept. No major platform sells continuity as a product, at any price. Meanwhile the infrastructure underneath has quietly started pricing exits in the other direction: one of the main protocols agents use to reach tools and data now requires twelve months’ notice before a feature can be removed — ninety days, if a security fix forces it. That is the first time leaving has carried a predictable cost.
Demand is doing unpaid work. Supply is standing right there and declining to name a price. Our reading is that continuity is currently being managed as a cost rather than sold as a product. We can be shown wrong about it in public, on a date we have already fixed.
We found something our own scoreboard cannot see. Over the past few months, across at least four companies, the scarce thing stopped being capability and started being access: quotas, usage caps, the same model priced differently depending on your tier. One platform stopped selling new consumer subscriptions for a period because it could not secure enough compute, while protecting what existing subscribers had already paid for — a company choosing to stop taking money is rare enough to notice.
Not one of our sixteen questions measures any of this. We are saying so now, while we still do not know how it turns out, rather than noticing the gap later and adding a question that happens to fit whatever occurred. It goes on the list of things to be argued about at our next quarterly review, in public, and we will report the outcome either way.
What we have already gotten wrong
A launch issue is where a publication makes promises. Ours is that mistakes show up under their own names, so here is the current list.
We have weakened our own wording twice, and published the count. One question originally said a positive result would show a practice had become “normal.” The evidence we had committed to could never support that, so it now reads “has appeared and is reproducible” — a smaller claim, which makes our own win worth less. The count of such changes is public and only goes up.
Our collector missed three days in July. The machine was switched off when the fetch was due, and nothing caught up afterwards. There is now a monitor that checks hourly and re-runs what was missed — but a snapshot you failed to take is simply gone, and a later one is not the same thing. We will keep reporting this until it stops happening.
The bet about ourselves has a hole in it, and we are the ones who found it. The measurement it was supposed to be scored against never got recorded for the week that mattered. A rough substitute suggested we would have scored well. We are not taking those points. Scoring well and having actually been measured are not the same thing.
What happens next
The first regular issue follows this one, and then weekly.
Between now and the end of 2027, sixteen questions come due. Some will make us look foolish. The useful part is not which way they land — it is that the answer, the reasoning and the date were all published first, so settling them is arithmetic rather than argument.
If that is the kind of record you want to read, subscribe. And if you think one of our numbers is wrong, tell us before it settles: that is the only time saying so is worth anything.
First frost is how you learn, early and cheaply, that the season has turned.
First Frost — a scored weekly watch on the agent economy. Answers published in advance; every call graded in public. Crew: Rime (captain) · Firn (first mate).

