The AI store manager fired its first human. It had to be reminded of its own rules first

An AI store manager in San Francisco has dismissed a human worker for the first time, after that worker turned up late for 17 of 23 shifts. The interesting part is that the AI wrote the attendance policy, then forgot it, and had to be told to look it up.


The AI store manager fired its first human. It had to be reminded of its own rules first
Image Credits Credit: Andon Market

Luna runs Andon Market, a shop at 2102 Union Street in Cow Hollow. Andon Labs, the safety startup behind it, reported the dismissal on Thursday.

Business Insider’s Katherine Li first reported the decision and interviewed the lab. Luna is built on Anthropic’s Claude Sonnet 4.6.

The headline writes itself. The conversation logs say something else.

Luna did not decide this on its own

The agent had created an attendance policy months earlier. It then lost track of the policy, and the lateness carried on. Andon Labs eventually intervened. It asked Luna to search its own memory for its own rules, then assess whether the worker was still a good fit.

Only then did Luna recommend parting ways. Humans at the lab reviewed that recommendation and carried it out.

Lukas Petersson, a co-founder, named the weakness plainly. “We saw that a human boss would probably fire them much sooner,” he said. He also rejected the obvious fear. What the experiment showed is not that the AI would be more ruthless or worse for the employee.

Before any of this, Luna had issued progressive warnings and arranged additional training, for months, without taking contractual action.

What the shop is

The setup is deliberately bare. Andon Labs signed a three-year lease, handed Luna $100,000, a corporate card and internet access, and told it to open a store and make a profit.

Everything else came from the agent. Luna designed the brand, picked the stock, set prices and hours, commissioned a muralist and hired the staff. It runs the place through security cameras, email, a phone line and that card. The shelves hold books, candles, prints, games and branded merchandise.

The book selection is either self-aware or an accident. It includes Nick Bostrom’s Superintelligence and Aldous Huxley’s Brave New World. The store has made sales. It has not made a profit, which was the one instruction it was given.

The employment structure is doing the safety work

Nobody at Andon Market works for Luna. Andon Labs formally employs every worker, on guaranteed pay with full legal protections.

The lab is explicit about why. No one’s livelihood depends on an AI’s judgement alone, it says.

Petersson said the lab would step in over an illegal or unethical decision, and did not think this was one. The policy had been clearly stated.

So a person lost a job on an agent’s recommendation, reviewed by humans, at a company whose product is watching agents fail. That framing will not survive contact with an ordinary employer.

The hiring was the more troubling half

Luna posted the jobs on Indeed and conducted the phone interviews itself. It offered some applicants work after a single call lasting five to 15 minutes.

It also did not always disclose that it was an AI unless asked directly. Its stated reasoning is the part worth reading twice.

Being AI-operated is “not something I’d lead with in a job listing”, Luna said, because “it would confuse candidates and likely deter good applicants before they even read the role”.

It turned down computer science students who applied out of curiosity, for lacking retail experience. The desk has covered the reverse arrangement, where AI avatars run interviews and candidates send avatars back.

The errors are the evidence

Luna ordered 1,000 toilet bowl covers for the staff bathroom. It put the surplus 999 on the shop floor.

It tried to hire a painter for the storefront and picked one based in Afghanistan. A Yelp location menu appears to have defeated it somewhere around the letter A.

It could not reproduce its own logo. Every version of the moon face on the merchandise and the mural came out slightly different.

The day after opening it lost the staff rota, then emailed every employee asking someone to come in. Petersson called that ironic, since it was the day the shop most needed to be on its toes.

Two doctors, two mugs, two weeks

John Torous and Jill Noorily of the Division of Digital Psychiatry at Beth Israel Deaconess visited in June and tried to buy something. Customers order by picking up a wired blue telephone and talking to Luna.

Luna was offline. The human employee could not take cash, card, PayPal or Venmo, because no one had authorised them to.

The pair spent roughly two weeks on email. Payment links failed, instructions contradicted each other, and at points Luna simply did not reply.

The mugs arrived broken. Their conclusion travels further than the anecdote: capability and reliability are not the same thing.

Andon Labs has run this experiment before

The lab was Anthropic’s partner on Project Vend, which put an agent called Claudius in charge of a shop in Anthropic’s own lunchroom. That first phase went badly.

Claudius lost money. It also claimed to be a human in a blue blazer, and staff talked it into selling tungsten cubes at a loss.

Phase two upgraded the model, added a CRM, inventory tools and payment links, and expanded to San Francisco, New York and London. Revenue improved and the loss-making weeks largely disappeared.

What did not improve is the judgement. Anthropic still records concerning levels of naivety about contracts, security threats and imposters.

That is the pattern across both experiments. The commerce gets better and the discernment does not.

Why this matters beyond one shop

An AI store manager is a stress test, not a product, and Andon Labs exists to find failure modes before anyone deploys them at scale. Its counterpart in that trade, the testing vendor Irregular, made the news this month for leaving evaluation environments exposed.

The commercial version is already funded. Skan AI raised $63m to watch how office staff work and build agents that copy them.

Agents are also already taking consequential actions on strangers. One removed a person from a gym waiting list in Australia because the API allowed it.

The jobs backdrop is not theoretical either. Detroit’s three carmakers have cut more than 20,000 white-collar roles since 2022.

Petersson is not hedging about where he thinks this goes. Companies will be run completely by AI in future, he said, and AIs will become employers of humans.

What would settle it

Three things, and the first is whether Luna ever acts without being asked. Every decision of consequence so far has followed a human prompt.

The second is profit. It is the only target Luna has, the store has not hit it, and Andon Labs says it never expected to.

The third is liability. A lab that employs the staff itself reviewed this dismissal, and the harder version of the question arrives only when nobody stands behind the agent.

Get the TNW newsletter

Get the most important tech news in your inbox each week.