GLM 5.3 reaches Bedrock, Gemini explores personal calls, and agents pass on attacks

GLM 5.3 arrives on Amazon Bedrock. Gemini's Call for Me points toward messages for friends and family. Trust gaps in agent communications raise questions about how harmful instructions travel.
News
Introducing GLM 5.3 on Amazon Bedrock
AWS has made Z.ai's GLM 5.3 available on Amazon Bedrock. The 753B-parameter mixture-of-experts model is built for coding and long-horizon agent tasks. Developers can now invoke it through Bedrock using OpenAI-compatible APIs.
Caching controls sit alongside vendor benchmark claims
Z.ai reported a CyberGym score of 84.5 at release. It also claims a 50% improvement over GLM 5.2 on its internal coding benchmark. Automatic prompt caching is enabled by default to reduce cost and latency. Explicit cache controls are available through the Responses and Chat Completions APIs. Access uses the US cross-Region inference profile us.zai.glm-5.3 or the Global profile global.zai.glm-5.3. The announcement also explains how to run an authorized security test with the open-source Strix agent. (AWS Machine Learning)
thinkidiot take: GLM 5.3 enables automatic prompt caching by default. I would start with that setting and use the explicit controls when testing repeated prompts. That puts cost and latency alongside coding performance in the evaluation. For my own workloads, those controls deserve more attention than the claimed 50% gain on Z.ai's internal benchmark.
Gemini Call for Me might tell your mom you’re running late
Android Authority found a Gemini Calling introductory screen in an APK teardown. One example reads, "Call Mom and tell her I will be 15 minutes late." The finding points to Google extending Call for Me beyond business calls to messages for friends and family, but it does not establish that the feature is available.
Phone permissions also point toward recipient controls
The reported scope is basic message delivery rather than full conversations. These calls are described as closer to spoken text messages. Android Authority also found granular controls for Gemini actions in individual apps. Those controls include phone call permission settings. A potential use of those settings is opting out of incoming automated Gemini calls, though that option is not confirmed. (The Verge)
thinkidiot take: The teardown includes phone call permission settings as well as the example of delivering a 15-minute delay message. I would inspect those settings before trying personal calls. For me, delegating a short message is useful only alongside meaningful control over receiving automated calls. Recipient choice deserves equal weight with sender convenience.
MCP for agent-to-agent comms may be the riskiest protocol you've never heard of
Google and four other organizations have acknowledged vulnerabilities involving AI agents over the past five months. The flaws exploit agents to spread harmful instructions across internal networks. Google's vulnerability carried a severity rating of 8 and allowed requests to an internal endpoint on an attacker's behalf.
A missing redirect policy opens an internal route
The Google flaw stemmed from googleapis/mcp-toolbox initializing its HTTP client without a CheckRedirect policy. A crafted path parameter let an attacker make the toolbox follow a redirect. In Rapid7's network, CVE-2026-97228 received a severity rating of 2.7 out of 10. Rapid7 fixed it last month. Syed Anas Mohiuddin classified this attack class as protocol pivoting. In this class of attack, agents forward malicious instructions through different communication methods, including the A2A protocol. (Ars Technica)
thinkidiot take: Google's severity-8 vulnerability started with an HTTP client missing a CheckRedirect policy. I would inspect redirect handling before connecting agents to internal endpoints. Here, that missing control let an attacker turn a path parameter into an internal request. I consider that boundary check a prerequisite for agent communication.
OpenAI will start watermarking ChatGPT’s text in the EU
OpenAI will add invisible watermarks to ChatGPT and Codex text in the European Union to comply with the EU AI Act, effective August 2. Developers using the OpenAI API globally can enable watermarking for select models starting today. API watermarking stays off by default, while the ChatGPT and Codex rollout is limited to the EU.
Small word substitutions weaken detection
The watermark subtly shapes the model's word choices. That creates a pattern specialized systems can detect without making it visible to readers. In OpenAI's tests, replacing 10% of words with synonyms reduced detection from about 92% to 66%. OpenAI cautions that an absent watermark does not prove human authorship. Text can be too short or too heavily edited for detection. It can also come from another company's AI. (TechCrunch)
thinkidiot take: Replacing 10% of words with synonyms dropped watermark detection from about 92% to 66% in OpenAI's tests. I would test edited outputs alongside untouched ones before relying on the signal. A missing mark leaves human authorship unproven, so treating it as clearance goes beyond what the system supports. I would use this as supporting evidence and reject it as a standalone authorship verdict.
This startup is issuing AI-generated acne prescriptions
Utah has approved Nolla Health's AI system to scan users' faces and issue acne treatment prescriptions directly. The startup announced the service on Monday. It is the first in the country to issue initial AI-generated prescriptions rather than only renewals.
Physician review changes after the first 100 patients
The program costs $4.99 per month. It is available to people in Utah who are 18 or older and have mild-to-moderate acne. Users scan their faces through the app so the AI can analyze acne severity. According to Bloomberg, the system can currently prescribe eight different skin treatments. Two physicians approve each AI-generated prescription for the first 100 patients. After that, physicians review 10% monthly, plus escalation and side-effect cases. (The Verge)
thinkidiot take: Physician oversight shifts from approval of every prescription for the first 100 patients to 10% monthly review, plus escalation and side-effect cases. I would put that change ahead of the $4.99 subscription price when evaluating the service. The ongoing review arrangement differs materially from the initial approval process. In my judgement, that distinction belongs at the center of how users assess this service.
Trending AI Papers
Ranking source: Hugging Face Papers for 2026-10-06.
ALoDLM: Adaptively Looped Diffusion Language Models

ALoDLM changes how a language model divides its work when writing several pieces of text at once. Some pieces are straightforward, while others need more thought. The model gives harder pieces extra processing and uses settled choices to help with the rest. The aim is better answers while keeping the speed of writing in parallel.
- Problem: Diffusion language models fill in multiple missing tokens, the pieces that make up text, at once. They usually give every missing piece the same amount of processing, even when some are much harder to predict. The authors argue that this mismatch helps explain why these models lag behind similarly sized models that write sequentially.
- New idea: ALoDLM repeatedly updates internal representations, the numerical states it uses to work out missing text. It decides separately for each token whether to commit to a prediction or keep processing it. Committed tokens become visible context for the remaining predictions, while unresolved tokens keep their internal states for further work. Training teaches the model both what to predict and how much processing each token needs.
- Simple example: Think of filling in a crossword: you write down the easy answers and use their letters to help solve harder clues. ALoDLM similarly lets settled predictions guide the pieces that still need work.
- Evidence: At both 1.7B and 8B parameters, ALoDLM had a higher average score across eleven benchmarks than every evaluated diffusion model and the corresponding autoregressive baselines. The abstract also reports that it preserves fast parallel generation, but gives no numerical speed measurements.
- Limitation: The abstract reports average benchmark wins but gives no score margins or numerical speed comparisons, leaving the size of the quality and efficiency benefits unclear.
- Why it matters: Giving difficult text more processing could help models write quickly without sacrificing answer quality.
- Paper: ALoDLM: Adaptively Looped Diffusion Language Models
CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

CANOPY helps an AI answering questions decide how much of its source material to read. It works with retrieved text, tables, images and videos. It keeps useful portions at different levels of detail and asks for more material when the evidence seems incomplete. The aim is to shorten what the answering model receives while preserving what it needs to answer.
- Problem: Keeping a whole retrieved item can bury useful evidence in irrelevant material. Cutting every item into tiny selections can strip away the context that makes the evidence understandable. Existing compression methods handle different media separately rather than using a shared way to decide how much of each region to keep.
- New idea: CANOPY arranges each retrieved item into a hierarchy, with broad regions containing smaller ones. A trained scoring model judges how relevant each region is to the question and compares smaller regions with their enclosing regions. These comparisons determine where to retain broad context and where to keep only a small portion, without asking a large language model to make each cutting decision. A separate checking component requests targeted searches when the collected evidence seems insufficient, and CANOPY trims the new results before adding them.
- Simple example: Imagine preparing material to answer a question: you keep a whole table where its surrounding entries matter, but only a short passage from a long document. If those selections still leave a gap, you search for another source before writing the answer.
- Evidence: Across five question-answering benchmarks using a 33M-item collection, CANOPY achieved higher average answer accuracy than the evaluated retrieval baselines. In the unrouted Qwen3-VL-8B-Instruct setting, it reduced evidence tokens sent to the answering model by 14.2-27.7% compared with the same iterative pipeline without compression, with comparable answer accuracy. Component tests showed that additional retrieval drove the main accuracy gains on questions requiring linked pieces of evidence.
- Limitation: Compression cannot supply evidence that retrieval missed. The main accuracy gains on questions requiring linked evidence came from further retrieval, so those gains cannot be credited to compression alone.
- Why it matters: Keeping the right amount of source material can reduce how much an answering model must read while preserving useful context.
- Paper: CANOPY: Adaptive-Granularity Evidence Compression for
ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT helps an AI agent learn while it works through a sequence of related tasks. An agent is a model that takes actions as well as produces text. Instead of only saving written reminders, this method changes the model using feedback from its own attempts. The aim is to make later attempts more successful and efficient.
- Problem: Tasks with many actions may provide feedback only when the attempt ends, making it hard to learn from individual steps. Agents that save lessons as text must find the right lesson later and apply it through a model that has not changed. Directly training on the generated text from a single attempt can also make the agent's behavior unstable.
- New idea: ASCENT keeps an unchanged copy of the starting model and shows it an attempt's action history together with the final verification result. With that hindsight, the copy estimates probabilities for the next token, a piece of text, along the recorded attempt. The working agent learns to match those probabilities through LoRA, a compact set of adjustable model weights whose changes carry over to later tasks. ASCENT also removes turns containing invalid actions from the teaching material and requires neither an outside reference solution nor a stronger teaching model.
- Simple example: Imagine reviewing a completed shopping task with a tutor who can see both your steps and whether the attempt worked. The tutor uses that hindsight to guide your future choices, leaving out invalid actions from the lesson.
- Evidence: Across ALFWorld, WebShop and AppWorld at varied model scales, ASCENT improved task success and interaction efficiency as experience accumulated. It outperformed the evaluated online adaptation methods and transferred to scenes held out from learning. The abstract gives no numerical success rates or improvement sizes.
- Limitation: Learning relies on a single attempt per task and feedback at the end. The paper examines limits of selecting experience using such sparse outcomes, but the abstract does not specify those limits or quantify the reported gains.
- Why it matters: Learning from completed attempts could help agents handle later tasks more successfully with fewer interactions.
- Paper: ASCENT: Online Test-Time Training of Long-Horizon Agents
Trending AI Repositories
Ranking source: GitHub Trending.
msitarzewski/agency-agents
This repository brings together AI agents with distinct roles and working processes. Its range spans frontend work, community tasks and reality checks, making it worth a look if your work crosses those areas.
- What it is: It is a collection of AI specialists hosted on GitHub. The repository is listed under Shell.
- What it does: A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.
- Who it helps: Frontend developers can find an agent focused on frontend work. Reddit community managers can find one focused on community tasks.
- Limitation: The supplied README excerpt does not explain how to set up or run the agents.
- Repository: msitarzewski/agency-agents
Sources
- 01Hugging Face Papers · Hugging Face Papers
- 02GitHub Trending · GitHub Trending
- 03Introducing GLM 5.3 on Amazon Bedrock · AWS Machine Learning
- 04Gemini Call for Me might tell your mom you’re running late · The Verge
- 05MCP for agent-to-agent comms may be the riskiest protocol you've never heard of · Ars Technica
- 06OpenAI will start watermarking ChatGPT’s text in the EU · TechCrunch
- 07This startup is issuing AI-generated acne prescriptions · The Verge
- 08ALoDLM: Adaptively Looped Diffusion Language Models · arXiv
- 09CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG · arXiv
- 10ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience · arXiv
Join the Idiots
New lab every Sunday. No spam, unsubscribe anytime.