OpenAI Ships a Cyber Model, a 14 MByte LLM Runs on Your Phone

OpenAI expands its cybersecurity program with a new cyber-trained model. Cactus Compute releases Needle2, an agentic LLM small enough for edge devices. nOps migrates to Bedrock AgentCore.
01.OpenAI Expands Daybreak and Ships a Cyber-Trained Model
OpenAI announced a major expansion of its cybersecurity program Daybreak, along with GPT-5.6-Cyber, a model trained specifically for offensive security research according to TechCrunch reporter Lucas Ropek on August 10. GPT-5.6-Cyber is available through Daybreak Red, OpenAI's authorized vulnerability research and exploit validation track, for approved cybersecurity professionals.
The company also published companion guidance on its own blog aimed at putting frontier cyber capabilities in trusted hands alongside the Daybreak expansion post.
Daybreak was originally launched as a bug bounty program. Its new iteration moves the model further into active defense operations, with approved partners gaining access to frontier capabilities for authorized vulnerability research and security testing. The move signals that OpenAI is treating its capability development as inseparable from its safety posture. An unusual strategy, since most companies split red teaming and blue teaming into different business units.
thinkidiot take: Pairing a bug bounty program with a custom cyber model is either confident or desperate. Either way, the industry needs more tools like this. The key metric here will be whether independent researchers can reproduce Daybreak's results at smaller models. If it only works on frontier-scale compute, it remains a luxury product.
02.Needle2: A 14 MByte Agentic LLM That Runs Anywhere
Cactus Compute released Needle2, an open agentic model at just 45 million parameters that runs as a 14 MB binary in roughly 28 MB of session RAM on its project page. The model targets phones, wearables, smart home hubs, and robots, where even the smallest quantized models have been too large for practical tool use.
Needle2 was trained for tool calling, device control actions, and structured extraction. Its training corpus covers smart home, mobile, wearable, TV, and car commands plus the BFCL benchmark suite. It achieved 61 overall on the full BFCL v4 single-turn test across 3,641 rows as reported on its own benchmark tables, with a well-formed rate of 93 across all rows. The gap concentrates where the model has never been trained: Java sits at 29, JavaScript at 32, versus 61 for Python, which aligns with its consumer device focus rather than general-purpose APIs.
The model is Apache 2.0 licensed hosted on GitHub. Researchers also published a companion paper at arXiv:2607.18363.

thinkidiot take: Forty-five million parameters sounds absurdly small to anyone who has read the last three years of AI headlines. That is exactly the point. The industry has spent so long treating parameter count as a proxy for quality that we forgot to ask what happens when you pick one specific job and do it well enough on cheap hardware. Needle2 does not win everywhere, but the market that matters here is the pocket device that cannot run anything else at all.
03.nOps Rebuilds FinOps Agents on Bedrock AgentCore in Four Months
nOps rebuilt its Clara FinOps AI agent on Amazon Bedrock AgentCore according to an AWS blog post earlier this week. The company replaced a self-managed Amazon EKS stack running LangChain and LangGraph, and cut its time to production from 10-12 months down to four.
The migration improved agent response quality and reduced operational overhead while keeping analytics governed through Databricks Lakehouse Metric Views per the same AWS case study. nOps kept its data pipeline in Databricks and swapped only the inference runtime, so Bedrock AgentCore handled the agent loop, tool calling, and orchestration.
thinkidiot take: The most interesting number here is not 75 percent. It is four months. A company building AI agents for over a year shipped its replacement in a quarter of the time by stopping the infrastructure management trade and using someone else's instead. That is a pattern to watch: the shift from building your own agent framework to plugging into managed platforms moves faster than most expect.
Sources
- 01As AI-led attacks multiply, OpenAI launches a new cyber model · TechCrunch
- 02OpenAI's Daybreak expands as the cyber defense window narrows · OpenAI
- 03Putting frontier cyber models in more trusted hands · OpenAI
- 04Needle 2 - The 14 MB Agentic LLM for Tiny Devices · Cactus Compute
- 05How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore · AWS Machine Learning
- 06Needle 2 release on GitHub · Cactus Compute GitHub
- 07A Controlled Study of Attention-Only Transformers · arXiv
Join the Idiots
New lab every Sunday. No spam, unsubscribe anytime.