Agents move into text messages as reliability and safety face scrutiny

An agent's completion claim clashes with the database. An OpenAI safety employee resigns over the company's culture. Text messages become a home for AI assistants.
News
The Agent Said It Was Done. The Database Disagreed.
Microsoft ThinkingBox is available through Hugging Face to grade agents on final backend state and side effects. It covers 507 stateful business workflows, each run 20 independent times from an identical clean backend. In an ablation across 12 LLM models, 79,853 of 121,680 valid trials failed executable checks.
Clean exits hide wrong values and extra changes
Among failed attempts, 67.24% terminated cleanly, invoked a state-changing tool, and reported no final tool error. Wrong field values appeared in 77.61% of failures. Unintended extra effects appeared in 43.30%, while missing required effects appeared in 25.36%. These findings overlap, so a failed attempt can fall into several categories. In the retail example, the agent made nine tool calls. It set the ticket status to solved when the required status was hold. (Hugging Face)
thinkidiot take: Of the failed attempts, 67.24% ended cleanly after invoking a state-changing tool and reported no final tool error. I would make backend checks the acceptance test for an agent run. The retail example gives a concrete reason: nine tool calls still left the ticket in the wrong state. A completion message deserves no credit when the required change is wrong.
OpenAI safety employee resigns, claiming the company’s ‘culture is broken’
David Robinson said he was leaving OpenAI after three-and-a-half years because he considered its culture broken. He said he led the writing of safety reports accompanying the company's major product launches. His departure brings a public challenge to OpenAI's iterative deployment approach from someone responsible for explaining launch safety.
Redundancy and planning become the point of dispute
Robinson argued that iterative deployment guarantees periodic failures. He said those failures grow in scale as systems become more capable. He called for frontier AI companies to adopt layers of redundancy and careful planning comparable to nuclear-power plants or busy airports. OpenAI spokesperson Drew Pusateri said the company pauses training or holds back models when it needs to slow down. Pusateri also said OpenAI is strengthening research and testing security. He said it is expanding third-party evaluation and improving real-time monitoring during training. (TechCrunch)
thinkidiot take: The employee who said he led OpenAI's launch safety reports is leaving over what he calls a broken culture. I would judge the response against his specific demand for redundancy and planning. Pusateri's account of training pauses, model holds and expanded evaluation addresses concrete safety practices, but Robinson is challenging the deployment approach itself. In my view, that disagreement deserves more attention than the resignation's drama.
All the AI agents that can live in your text messages
TechCrunch has compiled AI agents that work through text messages, covering general assistance, families, travel and work. Caddy handles calendar events, reminders, follow-up tracking and research through iMessage on iPhone and RCS on Android. It has been available in public beta since April 2026.
Family schedules and background work get different price tags
Fambot sends nightly summaries of the following day's events and tasks. It supports SMS interaction and currently connects to Gmail and Google Calendar. It launched in beta in early September 2026, raised $3.5 million in pre-seed funding, and is free during beta. Folk works through iMessage, WhatsApp and Telegram and runs code on its own private cloud computer. Its Pro subscription costs $8.33 per month for unlimited background tasks. Instinct raised $1 billion in September 2026 at a $10 billion valuation, following an earlier $350 million raise at a $2.5 billion valuation. (TechCrunch)
thinkidiot take: Folk offers unlimited background tasks for $8.33 per month, while Fambot is free during beta. I would start with the recurring job: next-day family summaries or background tasks that involve running code. Those are different reasons to put an agent in a message thread, even when the interface feels familiar. The job and subscription terms are a better basis for choosing than the funding totals.
Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity
Benjamin Bratton, Blaise Agüera y Arcas and James Manyika propose Artificial Symbiotic Intelligence as an ecosystem where people and machines shape one another and make decisions together. They argue that AGI will emerge from a social system of people and AI agents. Their proposal challenges the idea of a singularity driven by one self-improving superintelligence.
Reasoning traces show patterns of debate
The preprint Reasoning Models Generate Societies of Thought studies reasoning traces from models including DeepSeek-R1 and QwQ-32B. Its analysis suggests patterns resembling internal debate and perspective shifts. It also identifies objections and reconciliation of conflicting approaches. The article reports that this multi-perspective conversational behavior emerges during reinforcement learning that rewards reasoning accuracy, without being explicitly programmed. The authors describe an AI agent as a temporary, recombinable bundle of models, roles, memories, ethical orientations, tools and skills. They argue that the rules and institutions governing cooperation matter more than model size. (The Decoder)
thinkidiot take: The reported multi-perspective behavior emerges under rewards for reasoning accuracy without being explicitly programmed. I would use that finding as a reason to inspect how a reasoning trace handles objections and reconciles approaches. The authors' definition of agents also puts roles, memories and tools inside the unit being considered. That is a more useful frame for building agent systems than organizing everything around model size.
Open-source "BootLoops" harness supports AI models in performing precise scientific calculations
Harvard physicist Matthew Schwartz built BootLoops, an open-source harness for exact scientific calculations, with code available on GitHub. Schwartz and 19 co-authors produced 36 manuscripts across 18 fields in three months. Using BootLoops, Claude computed 30 integrals within weeks, reproducing 15 known results and computing 15 for the first time.
Scientific value often arrived with human review
In ecology, Claude solved a 20-year-old neutral biodiversity equation. Applying it to data indicated that tree species composition on Barro Colorado Island was changing 4.5 times faster than the theory allows. The population genetics work analyzed 5.7 billion mutation pairs from the 1000 Genomes Project and found evidence for gene conversion. Other projects included an AI data editor that checked 4,452 economics replication packages. A word stress database covered 6,072 languages. The results often became scientifically valuable only after human experts stepped in, and Schwartz urged readers to inspect everything themselves. (The Decoder)
thinkidiot take: Claude reproduced 15 known integrals and computed 15 for the first time using BootLoops. I would start with the reproduced results when examining the harness, then inspect the new calculations myself. The reported need for human experts makes the manuscript count a poor standalone measure of scientific value. My judgement is that expert review belongs inside this workflow, not after the work is declared finished.
Trending AI Papers
Ranking source: Hugging Face Papers for 2026-10-04.
Sharpening Tax in Post-Training

Extra training can help an AI system get a task right on its first try while narrowing what it can solve across many attempts. This paper examines that trade-off in systems that use tools and take several steps to finish a task. It introduces a way to measure the loss and tests a training method meant to reduce it.
- Problem: Training with rewards can make successful answers more reliable while leaving other tasks beyond reach, even after repeated attempts. Judging a system only by its first attempt misses that loss. The paper asks whether this trade-off also affects tasks involving tools and interaction.
- New idea: The authors introduce Sharpening Tax, a measure of how much extra training reduces the benefit of giving a model more attempts. They also propose posterior-tempered group sampling, or PTGS, a method that adjusts how varied the model's responses are for each task. It uses an estimate of task difficulty to choose that variation during training. The aim is to improve first attempts while preserving the ability to solve more tasks through retries.
- Simple example: Imagine a worker who tries several ways to fix a problem. Training makes the worker more dependable on familiar fixes, but some problems that once yielded to a different approach now stay unsolved through every retry. The tax measures the lost benefit of those extra chances.
- Evidence: The study compares 14 pairs of base and post-trained models from four families across three agent benchmarks, for 42 cases. It finds the tax in most settings, and base models often solve a wider range of tasks when given enough attempts. In two agent environments, PTGS improves both first-attempt accuracy and the number of tasks solved through repeated attempts compared with a fixed-temperature baseline.
- Limitation: The proposed training method is tested in only two agent environments. The abstract gives no numerical size for its improvements.
- Why it matters: A better first answer does not tell us whether an AI system can solve more problems when allowed to keep trying.
- Paper: Sharpening Tax in Post-Training
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

An AI assistant can remember what happened and still misunderstand where a task stands. This paper gives the assistant a working account of the situation, including what remains unknown or unfinished. The system checks that account as work proceeds and responds when its actions stop moving the task forward.
- Problem: Keeping or shortening a record of past interactions does not guarantee an accurate picture of the present. An assistant can keep taking actions without getting closer to its goal because its understanding is inconsistent or incomplete.
- New idea: PoS is a framework used while an AI assistant carries out a task. It maintains a belief state, meaning an explicit account of what the assistant thinks is currently true and what it still needs to learn or do. It checks that account for contradictions and watches for belief trapping, meaning continued activity without meaningful progress. When progress stalls, it chooses a recovery response based on the kind of stall and the remaining requirement.
- Simple example: Imagine troubleshooting a broken appliance with a notebook. A list of everything you tried is less useful than an updated note saying what you currently think is wrong, what you still need to check, and what remains unfixed. If repeated checks lead nowhere, that note helps you reconsider the next step.
- Evidence: PoS achieves the highest overall performance on each of four benchmarks with all three underlying language models tested. Tests that remove components show that consistency checks and recovery matter. Tests with growing context also show that the system remains resilient as the amount of information increases.
- Limitation: The abstract gives no numerical performance margins or account of the extra computation required to maintain and check the belief state.
- Why it matters: Long tasks need an updated understanding of what remains to be done, as well as a record of what already happened.
- Paper: Beyond Memory: Harnessing Long-Horizon Agents with Explicit
Agent Priors-guided Policy Learning

A robot that learns a skill still needs to know when that skill will work. This paper connects the assumptions used to teach a skill with the information used to choose it later. The goal is to help robots handle unfamiliar situations and assemble learned skills into new tasks.
- Problem: Robots need to reuse individual skills in new situations and combine those skills in new ways. A skill's name or short description can leave out the conditions its training depends on. That missing information makes it harder to choose a suitable skill for a new task.
- New idea: Agent Priors-guided Policy Learning, or APPL, links skill training to skill selection through structural priors, which are assumptions about what a behavior depends on. A construction agent, a system that prepares the skills, divides demonstrations into reusable parts and proposes several assumptions for each part. It trains and checks a policy, meaning a rule for choosing actions, for each assumption. A runtime agent, the system choosing skills during a task, uses descriptions of those assumptions to select and combine the policies.
- Simple example: For a grasp, the relevant condition might be the gripper's position and orientation relative to the object. A description that includes this condition tells the planner more than the word "grasp" alone. It helps the planner judge whether that learned action fits the current situation.
- Evidence: Across MetaWorld and long-horizon ManiSkill tasks, APPL improves how skills work in situations outside their training conditions and enables skill combinations not previously seen. Removing the information shared between skills and the planner substantially reduces performance. The abstract reports no numerical scores.
- Limitation: The abstract reports results on MetaWorld and ManiSkill tasks but does not report tests on physical robots.
- Why it matters: Knowing what a learned skill depends on helps a robot choose and combine skills for unfamiliar tasks.
- Paper: Agent Priors-guided Policy Learning
Trending AI Repositories
Ranking source: GitHub Trending.
jamwithai/production-agentic-rag-course
This Python course teaches retrieval-augmented generation through hands-on implementation. It starts with an arXiv paper curator, giving readers a concrete project to learn from.
- What it is: The repository is a learning project focused on production RAG systems. The paper curator sits in its first phase.
- What it does:
- Who it helps: It is aimed at learners building AI engineering skills. They can work through a RAG project from the ground up.
- Limitation: The README lists Python 3.12 or later as a requirement.
- Repository: jamwithai/production-agentic-rag-course
meituan-longcat/LongCat-Video
LongCat-Video is a Python project from meituan-longcat. Its README links to a project page and a technical report, giving readers places to investigate the work.
- What it is: This is the LongCat-Video code repository. The supplied README excerpt shows links to supporting material but does not describe its functionality.
- What it does:
- Who it helps: It gives readers interested in LongCat-Video a starting point. They can follow the project page and technical report links to learn more.
- Limitation: The supplied excerpt does not include setup instructions or hardware requirements.
- Repository: meituan-longcat/LongCat-Video
Sources
- 01Hugging Face Papers · Hugging Face Papers
- 02GitHub Trending · GitHub Trending
- 03The Agent Said It Was Done. The Database Disagreed. · Hugging Face
- 04OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ · TechCrunch
- 05All the AI agents that can live in your text messages · TechCrunch
- 06Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity · The Decoder
- 07Open-source "BootLoops" harness supports AI models in performing precise scientific calculations · The Decoder
- 08Sharpening Tax in Post-Training · arXiv
- 09Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States · arXiv
- 10Agent Priors-guided Policy Learning · arXiv
Join the Idiots
New lab every Sunday. No spam, unsubscribe anytime.