Agents pay per request, security scans go free, and Holmes’ desk becomes a website

BlockRun and Incarna use Amazon Bedrock AgentCore payments for pay-per-inference. Anthropic launches free security scans for open-source projects. A detailed website lets visitors explore a simulation of Elizabeth Holmes’ desk.
News
Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments
Incarna has put pay-per-inference into production using Amazon Bedrock AgentCore payments to pay BlockRun one request at a time. BlockRun serves more than 90 models from more than 15 providers over x402. AWS says the managed payment service cut Incarna’s x402 integration work from months to days.
Spending limits sit outside the model
AgentCore payments handles payment protocols, connects to wallets and signs transactions. Its infrastructure enforces spending limits outside the model. It supports x402-compatible endpoints, including Amazon Bedrock inference endpoints. Incarna provisions customer-owned agent wallets through the Coinbase CDP connector and uses delegated authorization. Payments settle in USDC on the Base network, where each transaction is verifiable on-chain. BlockRun quotes and settles each call independently. (AWS Machine Learning)
thinkidiot take: Spending limits are enforced outside the model. That is the part I would prioritize when connecting an agent to paid inference. I would set those limits before giving it delegated wallet access, then use the on-chain records to inspect its payments. For this setup, infrastructure-enforced budgets matter more than the size of the model catalog.
Anthropic launches free AI security scans for open-source projects
Anthropic has launched OSS Scanner for open-source projects. Projects that opt in receive security scans at no cost. The service runs those scans periodically.
The reports arrive without human triage
OSS Scanner uses Anthropic’s strongest models, including Claude Mythos, to generate reports entirely without human review or triage. Anthropic warns that the reports can be incorrect or invalid. (The Verge)
thinkidiot take: OSS Scanner delivers reports without human review or triage. I would treat each finding as a lead to reproduce before changing code. Free scanning still leaves me with the work of deciding whether a report is valid. I would use it as a source of investigation tasks, and give its findings no automatic authority.
Pretend you’re sitting at Elizabeth Holmes’ desk on this weirdly detailed website
Extend engineer Bo Lau has built a website that simulates Elizabeth Holmes’ desk. It draws on more than a thousand emails, slides, texts and documents unveiled during her trial. Visitors can explore that material through a recreation of her workspace.
A 2016 desktop doubles as a promotion
The simulation recreates a 2016 workspace with an iPhone with a home button and a MacBook Air running OS X El Capitan. Visitors browse public court documents through the simulated phone and laptop screens and can also run a simulated Theranos Edison machine. The project promotes Extend and its watch party for Nathan Fielder’s upcoming documentary, “You Can See Everything.” (TechCrunch)
thinkidiot take: More than a thousand trial records sit behind this simulated desk. I would open the documents before spending time with the Edison simulation. The period devices give the archive a setting, while the watch-party promotion gives the project a commercial purpose. I like this kind of interface when the documents themselves are the main attraction.
California is trying to shut down robot vs. human cage matches
The California State Athletic Commission has sent Rek a cease-and-desist letter over a human-versus-robot fight. Rek hosted the September 18 match between Frankie LaPenna and a humanoid robot. The commission called the fight unsanctioned.
A remote operator was behind the robot
A human controlled the robot through a remote virtual-reality system. Rek CEO Cix Liv posted the commission’s letter on September 30. The letter demanded commission approval for boxing or mixed martial arts contests involving any human in California. That demand also covered exhibitions. Rek’s next planned match on Friday features two robots fighting each other with weapons. (The Verge)
thinkidiot take: The commission’s letter demands approval for covered contests and exhibitions involving any human in California. I would make that approval part of planning a human-versus-robot match. The remote VR operator also matters to how I would describe the event: a person controlled the machine facing LaPenna. The robot format is interesting, but sanctioning belongs in the event plan.
Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect
Three fired OpenAI safety researchers have published an open letter disputing allegations that they mishandled sensitive information outside company procedures. Jasmine Wang, Tomek Korbak and Mikita Balesni deny those allegations. They say their dismissals have made colleagues afraid to raise concerns and collaborate with outside safety experts.
Questions about policies and protections went unanswered
The letter addressed OpenAI’s Safety and Security Committee, Safety Advisory Group and Mission Advisory Council. The researchers also denied involvement in a leak to The Information. That leak concerned architectures in OpenAI’s newest models that make chain-of-thought reasoning harder to monitor. OpenAI shared an internal memo denying retaliation. A spokesperson alleged an investigated pattern of misconduct beyond sharing information with an outside AI evaluation group. OpenAI did not directly answer TechCrunch’s questions about the specific policies allegedly violated or its protections for employees raising safety concerns. (TechCrunch)
thinkidiot take: OpenAI did not directly answer questions about the policies allegedly violated or protections for employees raising safety concerns. I would want those answers before accepting its misconduct explanation. The researchers’ denials and OpenAI’s allegations leave competing accounts, and the unanswered questions make those accounts harder to assess. Specific policy explanations are a reasonable standard for judging this dispute.
Trending AI Papers
Ranking source: Hugging Face Papers for 2026-10-09.
SuperNav: An Agentic Navigation System for Any Task in Any Scene

A service robot needs to work out what someone wants and how to get there. SuperNav puts a language model that can interpret images in charge of those decisions. Separate navigation tools handle movement. The aim is to help robots tackle different requests in places they have not encountered before.
- Problem: Training a model to choose robot movements can tie its abilities to the requests and places covered by its training examples. That leaves a gap when a robot faces an unfamiliar setting or a different kind of instruction.
- New idea: SuperNav uses a pretrained multimodal language model, a model that can interpret both words and images, without further training it specifically for navigation. It adds a support system with navigation instructions, tools that carry out movement, and records of the task's progress. The model marks a destination in an image to tell the movement tools where to go. It then uses feedback from those tools to reconsider its choices.
- Simple example: Think of a passenger pointing to a destination through a car window while a driver handles the steering. SuperNav makes a similar separation between choosing where to go and carrying out the movement.
- Evidence: SuperNav beat four evaluated baselines on tasks involving particular objects, multiple objects, and navigation guided by user needs. The authors also report category-level evaluation on HM3D and deployment on a real quadruped robot. The abstract gives no numerical performance scores.
- Limitation: The reported tests cover several task types and environments, but they do not establish that SuperNav can handle every request or scene.
- Why it matters: This work explores how robots can follow varied requests without training their decision model specifically for navigation.
- Paper: SuperNav: An Agentic Navigation System for Any Task in Any
From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

An AI agent needs somewhere to practise taking actions and responding to what happens. Trace2Env builds a stand-in for an unavailable system from records of past interactions. A language model uses those records to respond to the agent and keep track of changes. The goal is to make practice sessions behave like the original system over a sequence of actions.
- Problem: The systems needed to train or test an AI agent can be unavailable or too difficult to recreate. Existing language model simulations driven by prompts can struggle to reproduce responses and keep earlier actions consistent with later events.
- New idea: Trace2Env turns interaction traces, records of past actions and responses, into a reusable reference called a worldbook. This reference describes the environment's structure, evidence from the records, and inferred rules for how it behaves. A world model agent, a language model acting as the environment, consults that reference and a lasting record of the current interaction. It decides what each action reveals and what changes should persist, without additional model training.
- Simple example: Imagine rehearsing a conversation with someone who has studied records of earlier exchanges. They use those records to decide how to respond, while remembering what has already happened in your rehearsal. Trace2Env applies that idea to an agent interacting with a simulated system.
- Evidence: Across nine environments, Trace2Env reproduced the next observation more faithfully and maintained consistency over longer interactions better than conventional prompt-based language world models. Actions produced during its simulations were also valid more often when replayed in the real environment. The abstract gives no numerical size for these improvements.
- Limitation: Trace2Env depends on access to historical interaction records. The abstract does not establish how well it handles behavior those records do not cover.
- Why it matters: It offers a way to train and test agents when the original system cannot be used or rebuilt.
- Paper: From Traces to Agentic Worlds: Agentic Language World
TokenRouter: Efficient Serving System for Token-Level LLM Routing

An AI answer can draw on different language models as it is being written. Switching models for individual pieces of text creates scheduling problems for the software running them. TokenRouter is a system designed to manage those switches. It aims to increase text generation speed and make the switching rules easier for developers to write.
- Problem: Token-level routing means choosing a model for each token, a small piece of generated text. Systems built around a single model struggle when those choices make generation steps fall out of sync and requests wait to join processing groups. Developers also face complicated implementation work.
- New idea: TokenRouter lets developers write routing rules, which choose the next model, from the perspective of an individual request. The system gives each model its own serving process and sends requests to it without requiring all models to move together. Each process uses delayed batching, which briefly postpones grouping requests for joint processing. A mathematical model of processing capacity determines the settings for that scheduling method.
- Simple example: Think of a job moving between specialist desks. The instructions describe which desk the job needs next, while each desk gathers work it can handle together and proceeds on its own schedule.
- Evidence: Across different routing algorithms, workloads, and model pairs, TokenRouter achieved 2.01 to 64.15 times the decoding throughput of existing systems. Decoding throughput measures how much generated text a system can produce per unit of time.
- Limitation: The abstract reports throughput gains but gives no measurements of how long an individual request takes to finish or how much serving costs.
- Why it matters: TokenRouter makes switching models during text generation more efficient to run and simpler to implement.
- Paper: TokenRouter: Efficient Serving System for Token-Level LLM
Trending AI Repositories
Ranking source: GitHub Trending.
anthropics/knowledge-work-plugins
Knowledge Work Plugins gives work preferences a place in Claude's setup. The useful idea is that setting a goal can also mean specifying how the work should be done.
- What it is: These plugins bundle working preferences for Claude Cowork and are compatible with Claude Code. Users can specify which tools and data Claude uses, how it handles workflows, and which slash commands it exposes.
- What it does: Anthropic's open-source plugin collection for Claude Cowork (also compatible with Claude Code) that turns Claude into a specialist for a specific role, team, or company. Each plugin bundles the skills, connectors, slash commands, and sub-agents for a job function (productivity, sales, marketing, data analysis, and more) so you can steer which tools and data Claude pulls from and how it handles critical workflows.
- Who it helps: It helps people who want Claude to follow the working practices of their role, team, or company. They can specify which tools and data it uses, how it handles workflows, and which slash commands it exposes.
- Limitation: The stated compatibility covers Claude Cowork and Claude Code.
- Repository: anthropics/knowledge-work-plugins
Sources
- 01Hugging Face Papers · Hugging Face Papers
- 02GitHub Trending · GitHub Trending
- 03Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments · AWS Machine Learning
- 04Anthropic launches free AI security scans for open-source projects · The Verge
- 05Pretend you’re sitting at Elizabeth Holmes’ desk on this weirdly detailed website · TechCrunch
- 06California is trying to shut down robot vs. human cage matches · The Verge
- 07Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect · TechCrunch
- 08SuperNav: An Agentic Navigation System for Any Task in Any Scene · arXiv
- 09From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation · arXiv
- 10TokenRouter: Efficient Serving System for Token-Level LLM Routing · arXiv
Join the Idiots
New lab every Sunday. No spam, unsubscribe anytime.