Claude Haiku 5.5 on AWS, edge decision models and games from text

Claude Haiku 5.5 arrives on AWS. Open d1 models bring multimodal decisions to the edge. Playground introduces a place to create and play custom games.
News
Introducing Claude Haiku 5.5 on AWS
Anthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. Anthropic calls it the fastest and most efficient model in the Claude 5.5 family. It says the model costs around 75 percent less than Claude Haiku 4.5 for most tasks. It is also the first Haiku model with effort controls, letting users tune cost against intelligence for each task.
A subagent for the work that piles up
Haiku 5.5 is built for subagents and high-volume, cost-sensitive work. In coding workflows, it can route requests and review code. It also classifies long documents. Its capabilities include agentic coding and multi-step tool use. The model supports high-resolution images. (AWS Machine Learning)
thinkidiot take: Around 75 percent lower cost than Haiku 4.5 for most tasks is a concrete reason to revisit routine workloads. I would start with request routing and document classification, then adjust effort for each task. Those controls let me make separate cost decisions inside the same workload. For this kind of work, I value that control more than the claim of being the family's fastest model.
Multimodal open d1 decision models for the edge
An announcement on Hugging Face introduces the multimodal open d1 decision models for the edge. The d1-3B model scores 48.57 on Decision Index 0.2.1. That puts it ahead of every 4B and 9B model in the reported comparison. It also beats Decider 35B-A3B, which scores 47.11.
Question latency reaches 16 ms on Jetson hardware
The reported time for d1-3B to answer a question on an NVIDIA Jetson AGX Thor is 16 ms. On a Jetson AGX Orin, it takes 26 ms. The listed result for a Jetson Orin Nano is 50 ms. The announcement also reports a mean score of 82.9 for d1-3B, the highest in its table and above Decider 4B. The smaller d1-omni-600M scores 78.4, ahead of Decider 2B at 77.1. The announcement describes that result as using only a quarter of the parameters. (Hugging Face)
thinkidiot take: The d1-3B model beats Decider 35B-A3B on Decision Index 0.2.1, scoring 48.57 against 47.11. I would use that result to put the smaller model first in an edge evaluation. The listed Jetson timings give me specific hardware results to test against. For choosing what to try first, I find that combination more persuasive than parameter count.
Introducing Playground: Create and play custom games
Google has introduced Playground, an experimental platform for custom games. It lets people create games with simple text prompts and no coding experience. Users can play their creations and share them with other people.
Publishing brings ratings, discovery and safety screening
A finished game can stay private. A shareable link lets friends and family play, while publishing puts it in the Playground Explore gallery. With a Play Games profile, users can claim a custom handle and like their favorite games. They can also compete for the top spot and follow creators for new releases. The gallery uses player ratings and play activity to spotlight fresh, creative and fun games. Every published game undergoes safety screenings aligned with Community Guidelines. (Google)
thinkidiot take: Playground lets someone with no coding experience turn text prompts into a playable game. I would start with a private game and send a link to friends before publishing it. That sequence gives me a way to share a creation before putting it into a gallery shaped by ratings and play activity. The choice over who gets to play is the feature I value most here.
Nous Research confirms it hit $1.5B valuation, launches AI agents for business users
Nous Research has confirmed a $90 million Series B at a $1.5 billion valuation, matching TechCrunch's earlier reporting. Robot Ventures led the round. The developer of Hermes Agent is using the capital to push into the enterprise sector with Hermes for Businesses. The offering lets companies deploy customized agents for multi-step workflows, with Nous promising to keep their data private and secure.
Revenue targets follow a large open source footprint
Nvidia, Union Square Ventures, Menlo Ventures, Samsung and 1789 Capital also participated in the round. Donald Trump Jr. is a partner at 1789 Capital. According to Nous, the open source Hermes Agent has been cloned more than 24 million times. The startup estimates that it drives roughly 2.5 percent of global AI token usage. Nous was at roughly $36 million in annualized revenue by mid-September 2026. It expects to pass $100 million before the end of 2026. (TechCrunch)
thinkidiot take: Nous is taking Hermes into business workflows after reporting more than 24 million clones of its open source agent. I would evaluate the business offering on a concrete multi-step workflow and its handling of private data. Clone counts and estimated token usage do not answer those deployment questions. For my purposes, the enterprise pitch earns its place through workflow execution and data controls, not the $1.5 billion valuation.
Microsoft releases new Nvidia-chip AI PCs with revamped Windows 11
Microsoft has revealed specs and prices for the Surface Laptop Ultra, using Nvidia chips designed to run AI models and agents. The base models start at $2,600 and $3,700, with the latter using a more powerful chip. Prices reach $5,900. Microsoft also revealed the Surface RTX Spark Dev Box, a workstation starting at $6,000.
Agent sandboxing extends to all Windows 11 users
Both devices are designed to run AI models locally for free, using their CPU, GPU, unified memory and cooling. The Dev Box uses the RTX Spark chip. It comes equipped with VS Code, GitHub Copilot CLI, WSL and PowerShell 7. The revamped Windows 11 includes Execution Containers to make sandboxing AI agents easier. That feature will be available to all Windows 11 users. Microsoft is also offering up to $1,000 off for a MacBook Pro trade-in, as Macs have become popular with developers and AI agent enthusiasts, particularly in Silicon Valley. (TechCrunch)
thinkidiot take: Execution Containers will be available to all Windows 11 users, while the new laptops start at $2,600. I would evaluate that sandboxing feature before buying hardware for local agents. Free local model execution still comes with a substantial purchase price on these machines. For me, the broader Windows feature is the more consequential part of this announcement.
Trending AI Papers
Ranking source: Hugging Face Papers for 2026-10-08.
GRACE: Generation-aware latent compression for efficient video generation

GRACE aims to make video generation faster by shrinking the information a model must process. Shrinking that information usually hurts image quality or forces costly changes to an already trained model. This work changes how the compressed information is learned so the existing generator can still use it.
- Problem: Stronger compression loses detail, while adding capacity to recover that detail slows the generator's training. Training compression only to reproduce the original video also changes the information the generator receives, making it harder to reuse a pretrained model.
- New idea: GRACE keeps a fixed base latent, an internal representation produced by the original video encoder. It learns a residual latent, a separate representation of information lost through stronger compression. It also aligns the compressed representation using features from the fixed Diffusion Transformer, the model that turns noisy representations into video. Finally, it gives that generator a small amount of additional training and teaches it to remove noise from the base before the residual.
- Simple example: Think of packing a detailed drawing as a familiar outline plus a smaller set of corrections. The outline preserves the structure, and the corrections restore missing detail. GRACE likewise separates a base representation from the information needed to fill its gaps.
- Evidence: On Wan2.1-I2V-14B at 480x832x81, GRACE reports 8x fewer tokens, the units the generator processes, and an 11.1x reduction in latency. It matches the original pipeline's generation quality on VBench.
- Limitation: The abstract reports results for one named video model at one output size. It does not establish whether the same speed and quality results hold for other models or sizes.
- Why it matters: It offers a way to speed up video generation while reusing a trained generator and preserving measured quality.
- Paper: GRACE: Generation-aware latent compression for efficient
Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

An AI model can make its own practice questions, but that practice can eventually make it worse. This paper traces the decline to flawed questions and repeated mathematical tasks disguised by different wording. R-Quest changes how questions are judged so models can keep learning from their own exercises.
- Problem: Invalid questions become more common over repeated training rounds, and selecting questions by agreement among answers increases their share further. Checks based on wording also miss questions that look different but ask the same mathematical thing, leaving the model with increasingly repetitive practice.
- New idea: R-Quest trains the solver, the model that answers questions, to identify and reject invalid questions. Those judgments shape rewards for the questioner, the model that writes practice questions, and determine which questions enter the solver's training data. A frozen base model, a reference model whose training stays fixed, compares sampled pairs of questions. Its judgments provide novelty feedback, a signal about whether questions offer different practice.
- Simple example: Imagine a student writing their own practice sheet. Rewording the same exercise fills the page without adding variety, while flawed questions make the practice less useful. R-Quest adds checks for both problems before the questions become training material.
- Evidence: R-Quest achieves the highest average performance across 12 benchmarks covering mathematical reasoning, general-domain reasoning and code generation, with tests across two model families. It sustains gains over ten rounds of self-training and outperforms R-Zero by 17.32 points.
- Limitation: The abstract tests sustained improvement over ten rounds. It does not establish whether the gains continue beyond that point.
- Why it matters: Better checks on self-written questions can help a model keep improving instead of learning from flawed or repetitive practice.
- Paper: Questioning the Questions: Sustaining Self-Evolution in
nanoMuse: An Open-Source Personal Agent for Every Device You Own

nanoMuse proposes an open-source assistant that stays with a person across their devices. It is designed to keep a continuing conversation, remember information and act through phone and computer screens. The report describes the software and sets out work still needed to develop and evaluate it.
- Problem: Earlier assistants waited for requests, while later agents completed a task and stopped. The report describes Meta's Muse as a continuing personal agent, but one confined to a vendor's cloud and one country, with no open counterpart until now.
- New idea: nanoMuse connects a person's devices through a relay, a service that carries their shared conversation and that anyone can run. It can operate phone and computer screens, with every action passing through a component called Sentinel. Its memory is stored in files the person can read, and the person chooses the AI model it uses. The software is released under GPL-3.0, an open-source license.
- Simple example: Think of one shared notebook that follows you between your phone and computer. nanoMuse applies that idea to an assistant's conversation, while also giving it ways to act on both screens.
- Evidence: The abstract presents an open-source implementation under GPL-3.0 and an analysis of Muse based on Meta's public record and a copy of its production prompt. It reports no numerical evaluation results, and describes size and cost figures as estimates.
- Limitation: The evaluation suite for screen actions remains on the roadmap, alongside memory that records where information came from and an open model for screen actions. The abstract therefore leaves the reliability of those actions untested.
- Why it matters: It gives people an open-source route to a continuing personal assistant with readable memory and a model they choose.
- Paper: nanoMuse: An Open-Source Personal Agent for Every Device
Trending AI Repositories
Ranking source: GitHub Trending.
manaflow-ai/cmux
cmux puts AI coding agent work at the center of a terminal app. Its focus on keeping that work organized makes it worth a look if you juggle several agents.
- What it is: This Swift project is available as a downloadable macOS app.
- What it does: Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.
- Who it helps: It helps developers who work with multiple AI coding agents. They can organize sessions in vertical tabs and receive notifications.
- Limitation: It requires macOS.
- Repository: manaflow-ai/cmux
Sources
- 01Hugging Face Papers · Hugging Face Papers
- 02GitHub Trending · GitHub Trending
- 03Introducing Claude Haiku 5.5 on AWS · AWS Machine Learning
- 04Multimodal open d1 decision models for the edge · Hugging Face
- 05Introducing Playground: Create and play custom games · Google
- 06Nous Research confirms it hit $1.5B valuation, launches AI agents for business users · TechCrunch
- 07Microsoft releases new Nvidia-chip AI PCs with revamped Windows 11 · TechCrunch
- 08GRACE: Generation-aware latent compression for efficient video generation · arXiv
- 09Questioning the Questions: Sustaining Self-Evolution in Reasoning Models · arXiv
- 10nanoMuse: An Open-Source Personal Agent for Every Device You Own · arXiv
Join the Idiots
New lab every Sunday. No spam, unsubscribe anytime.