xAI ties GPT-5.6 in frontier benchmark as Microsoft moves toward a Copilot super app

xAI's Grok 4.6 scores 61 on the AI Intelligence Index alongside OpenAI's best, while Microsoft merges its consumer and commercial Copilot apps into one interface.
01.xAI releases Grok 4.6, matching GPT-5.6 in Artificial Analysis benchmark
xAI has released Grok 4.6, which scores 61 points on the Artificial Analysis Intelligence Index, tying with OpenAI's GPT-5.6 Sol at that level. The model trails only Anthropic's Claude Opus 5 (which leads the index) and costs more than 60 percent less on agentic tasks, xAI says.
The comparison is measured through a task-based intelligence index from Artificial Analysis, which scores models across capabilities that go beyond standard benchmark scores. On agentic workflows specifically, Grok 4.6 completes complex multi-step tasks in about 53 steps where Claude Opus 5 averages 103, according to the Decoder's report of the findings.
xAI announced Grok 4.6 on The Decoder. The company claims it was built using data from "open research papers, books, and public knowledge repositories" rather than its own training data scraped from X.
thinkidiot take: A model scoring 61 out of roughly 90 on a capabilities-index is not the end state of frontier AI, but the gap between 53 agentic steps and 103 for Claude Opus 5 is meaningful in practice if you actually run workflows rather than answer multiple-choice questions. Whether "more than 60 percent lower" pricing holds up at scale matters more than index scores.
02.Microsoft combines Copilot apps ahead of a super app launch later this year
Microsoft has started merging its consumer and commercial Copilot AI assistants into a single interface, beginning with the standalone Copilot app and the Microsoft 365 Copilot app. Both personal and work accounts will route to the same unified application under the recycled "Microsoft Copilot" name but with an updated icon.
The change removes one of the two redundant Copilot taskbar icons that many Windows users had complained about. A global rollout for mobile and web apps starts in mid-August, according to Microsoft's update announcement, with Windows and Mac app updates arriving by mid-September.
Podcasts and Deep Research are being retired from the unified experience and will no longer be available after August 18th. Group Chat threads, messages, and content will also not carry over to the new app. Core Copilot capabilities remain free, though Microsoft notes that some users may see different usage limits for the free tier under the restructured app.
Microsoft describes the changes on The Verge. Mobile app users will be prompted to download the unified version, and some Windows Insiders should see the update this week.
thinkidiot take: Consolidation makes sense when you have two copies of the same assistant cluttering your desktop, but retiring Deep Research is a hard pill. The "super app" Microsoft envisions later this year will likely be a commercial play disguised as convenience. That does not mean it must fail, but history suggests users do not want a single app for everything they do on their phone or computer.
03.Apple considers paywall-breaking deal to get current news into Siri
Apple is in talks with publishers about paying nine figures to provide Siri with current news content, according to an article on TechCrunch that cited the Wall Street Journal. The move would represent a significant departure from Apple's traditional approach of keeping Siri focused on its walled-garden ecosystem.
The full report appeared on TechCrunch, describing a nine-figure budget for the payments. Apple has long struggled to make Siri feel useful in ways that compete with Google's voice assistant, particularly when news and information queries are involved.
thinkidiot take: This is a reasonable bet for Apple. If Siri reads today's headlines aloud on your commute, it becomes marginally less irrelevant compared to Google Assistant. But paying publishers nine figures is not the same as building better AI, and I would wait to see what kind of coverage you actually get before buying into the narrative.
04.Anthropic introduces a benchmark for reasoning about truths you cannot easily verify
Anthropic published the Conceptual Reasoning Index (CRI), a new benchmark targeting questions whose answers are "practically impossible to verify empirically or mathematically." Traditional benchmarks measure what a model knows. CRI measures whether a model can reason coherently when standard truth-checking does not work.
The paper is available on Anthropic's alignment research blog. This is the kind of evaluation tool that becomes more important as AI systems are deployed for decisions where every claim cannot be fact-checked against a source.
thinkidiot take: A benchmark for "what you cannot prove" is exactly what you need when frontier models are no longer primarily tested on math problems that have single correct answers, but deployed in contexts where the stakes of being wrong remain unclear. This is not peer reviewed, and the index likely reflects Anthropic's own priorities more than a neutral ground truth, which is normal for lab-designed benchmarks.
Five ML papers worth reading
-
AVA-Encoder: Towards Agent-Native Video Representation Learning (arXiv:2608.12313) -- Creative agents lack a structured video representation for learning from high-quality human films. This paper proposes an encoder specifically designed so that agent-native video embeddings respect film grammar rather than treating frames as isolated snapshots. Relevant for anyone building generative video systems beyond simple clip prediction.
-
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation (arXiv:2608.12308) -- Introduces a diffusion-planning approach combined with causal memory for agents navigating 3D environments from visual-language instructions. A step toward robust drone and mobile robot navigation where the agent must follow natural language commands in unfamiliar settings.
-
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307) -- Explores transferring capabilities from strong models to weak ones at inference time using harness-based interventions, rather than through training or fine-tuning. Useful if you have a large model's output but need to run cheaper, smaller alternatives in production.
-
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations (arXiv:2608.12304) -- Reviews how class activation mapping has evolved from early CNN-based visual explanations through transformers to foundation-model-era techniques. A comprehensive overview for researchers needing to understand interpretability methods across model generations.
-
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies (arXiv:2608.12307) -- Proposes a framework for evaluating multi-hop reasoning in agents that must use API tools under specific policy constraints, a gap in current evaluation tooling as agentic systems grow more complex.
Five AI repositories worth exploring
- mcp-memory -- Fast agent memory using Google's OKF and SQLite FTS5, built for the MCP (Model Context Protocol) ecosystem. Provides an efficient storage layer for agents that need persistent cross-session context without heavy infrastructure.
Additional papers from the previous 24 hours in arXiv
-
Redistribution-based Cost Inference Improves Sparse Safe Offline RL (arXiv:2608.12306) -- Reformulates cost inference in offline reinforcement learning through redistribution rather than direct estimation, enabling sparse-cost safe RL with polynomial-time guarantees on reward term selection via min-cover on causal DAGs.
-
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence (arXiv:2608.12302) -- Frames the image-to-video generation problem as an agentic optimization task that actively aligns generated video with input imagery rather than passively training through trial and error.
Sources
- 01xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index · The Decoder
- 02Microsoft is combining its Copilot apps ahead of a super app · The Verge
- 03Apple in talks to pay publishers to provide Siri with current news · TechCrunch
- 04Introducing the Conceptual Reasoning Index · Anthropic
- 05AVA-Encoder: Towards Agent-Native Video Representation Learning · arXiv
- 06DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation · arXiv
- 07mcp-memory on GitHub · GitHub
Join the Idiots
New lab every Sunday. No spam, unsubscribe anytime.