> TECHZIP TODAY | TLDR

Good morning, and welcome to today’s Techzip.

The biggest shift is that AI agents are moving beyond one-off prompts: OpenAI is testing a persistent Codex, Thinking Machines is training models with task expertise, and Anthropic is pushing agents toward physical equipment.

Add a warning from 100+ organizations about AI-enabled cyberattacks, and the message is clear: capability is expanding fast, but memory, access, and control are becoming the real product questions.

> SPOTLIGHT

WHAT MATTERS TODAY

WIRED reports that OpenAI is testing a Persistent mode in the Codex CLI. The agent is designed to keep working until it is “put to sleep,” reuse prior interactions and knowledge, create follow-up tasks, and message the user selectively. OpenAI has confirmed the testing, but has not announced a launch date.

The shift is not simply that an agent can run for longer. It turns an agent from a tool that completes a prompt into an ongoing operating layer: one that remembers context across sessions, notices unfinished work, and returns to it. But persistence is not autonomy. The agent does not gain new authority on its own, and actions outside its allowed scope still require approval.

That makes the sleep button, memory boundary, and approval flow part of the product lifecycle, not minor settings. A persistent agent becomes genuinely useful only when a team can answer four questions: What may it remember? What may it follow up on? When must it ask? And who can put it to sleep?

Thinking Machines Lab audited 2,500 examples from BIRD and says 61.1% contained at least one error, spanning the question, the SQL answer, or the supporting knowledge. Rather than simply giving a model more tools and asking it to work things out, the team had experts review the data, define what a correct answer looks like, and use those standards to further train Kimi-K2.6.

The lab reports 91.37% under its default run and 92.97% when sampling multiple answers and selecting the best result. Those are company-reported results, not proof that Text-to-SQL is solved. Still, they point to something useful: for a narrow enough task, teaching an AI the standards of strong practitioners can work better than making it infer those standards while doing the job.

The practical point is that expertise does not live only in prompts or final review. If a team knows which mistakes recur and what a correct result looks like, it can put those standards into training data and help the system get more right from the start.

More than 100 organizations, including OpenAI and Anthropic, signed a warning that companies and critical infrastructure may have only months to prepare for AI-enabled cyberattacks. It is a collective warning, not evidence of a confirmed attack timeline.

But the concern is not limited to novel vulnerabilities: excessive permissions, misconfigurations, unpatched software, weak authentication, and legacy systems can all become cheaper entry points when attackers have AI help.

The story is therefore not only about defending against AI. It is also about using AI for defense before attackers scale. The signatories call this the “defender’s window”: companies raise their security bar, security vendors continuously test defenses and share validated threat intelligence and playbooks, and governments coordinate support for under-resourced infrastructure.

> PRESENTED BY TOP PROVIDER

Find The Right Payment Processing Providers For Your Business. ​In Minutes, Not Months.

Finding the right payment processor shouldn’t take weeks of wrong vendors and inbox spam. There’s a faster way with Top Provider. Answer a few questions about your business: your industry, your volume, and how you take payments. In minutes, we match you with processors that fit your specific situation. No research rabbit holes. No chasing quotes. The right fit, faster than you thought possible. Get started with Top Provider today.

> SIGNAL HEADLINES

CAPTURE THE SHIFT

  • Perplexity Brain Adds a Memory Layer to Perplexity Computer: Perplexity says Brain turns sessions, files, and sources into a knowledge wiki maintained by background agents. Its correctness, currentness, and recall gains are company-reported, but the direction is clear: memory is becoming a competitive layer in agent products.

  • Gemini Omni 1.1 Flash Moves Video AI Into an Editing Loop: GoogleAI says the model adds video generation and editing, 4K upscaling, first-and-last-frame controls, and scene extension. The workflow is moving beyond one-shot clips toward creating, revising, and connecting scenes with context.

  • Qwen3.8-27B Gets a Broader Local Deployment Ecosystem: QwenDevs highlights GGUF, MLX, AWQ, NVFP4, and multiple community serving options. With open models, practical value increasingly depends on how a model is packaged, quantized, and deployed, not only on the base release.

  • Hunyuan 4 Preview Comes With a Leaderboard Reminder: Tencent describes Hunyuan 4 preview as 770B total parameters, 49B active parameters, and a 1M-token context window. Arena notes that early AutoEval results still need to converge with live human votes, making this a model signal to watch rather than a final ranking.

  • Jensen Huang Says “AGI” for Many Tasks, but the Benchmark Still Does Not Exist: Jensen Huang said that “for many tasks” we could call AI systems AGI, while The Verge notes that AGI has no agreed definition or benchmark. The useful reading is a claim about task capability, not a universally measured milestone.

  • Claude in Chrome Reaches Paid Plans: Anthropic has made Claude in Chrome generally available with a safety classifier. As agents that operate across the web become mainstream product surfaces, permission boundaries and prompt-injection defenses become part of adoption itself.

  • ChatGPT Work Adds Authentication for Supported Websites: OpenAI has added authentication for supported websites in ChatGPT Work. Once an agent has an authenticated session, the important question shifts from how good the prompt is to what access is granted, how long it persists, and which actions require approval again.

  • Anthropic Extends Agents to Physical Equipment With MHS: Anthropic has opened a research preview of Model Hardware Standard, intended to help agents discover, understand, and operate equipment in scientific research and manufacturing. The integration results are company-reported, but the action surface now reaches beyond the web and brings a new safety-interface problem with it.

  • Anthropic and MatX Keep the Custom-Chip Option Open: Reuters reports that Anthropic discussed buying MatX for about $7 billion before talks shifted toward a possible partnership. No transaction has been confirmed, but the story shows why frontier labs may want to preserve custom silicon as a strategic option.

> ONE PRACTICAL USE OF AI TODAY

Using Grok Bot to Design Grok Bot

Today, I bring to you this guide from SpaceXAI designer - John Bai.

He describes using Grok Bot as a team of AI design agents to design Grok Bot. The point is not to replace the designer. It is to prototype faster, explore more directions, and reduce repetitive production work while the human still decides which version is actually good enough to keep.

The design loop has four parts:

  • Assign agents by work type. Experiments turns vague ideas into prototypes. Motion God explores animation and interaction. Figma Bro handles production work in Figma. Devbot answers engineering questions. The bots can coordinate and delegate instead of forcing one agent to do everything.

  • Make an idea real enough to judge. When exploring how Grok Bot might live on a Mac, Experiments tried three directions: inside the notch, peeking from a desktop corner, or following the cursor. The idea did not need to be proven feasible or roadmap-worthy before it could be experienced.

  • Work on real production assets. Motion God used the actual animation asset and spec to build a local playground. The designer could adjust timing, spring, scale, and easing, or simply say “the bounce is too strong,” then watch the revised prototype immediately.

  • Automate repetition, not judgment. Figma Bro reads the actual components, spacing, typography, and file structure to populate cards, logos, and assets. It removes attention-heavy production work rather than taking over design decisions.

The loop is: idea → agent prototype → designer response → agent revision → designer judgment → repeat. The real value is not only speed. It is making the cost of trying an idea low enough to explore more directions before committing to one.

> WORTH READING

ANALYSIS & THESIS

A Tsinghua study covering 1,338 post-training runs finds that agents can train, debug, evaluate, and iterate, but rarely reconsider strategy. It is a useful counterweight to the idea that simply giving an agent more tokens will make it find its own way.

This essay argues that the AI debate should focus on nonproliferation and control, not only data centers. It reframes the compute race around a harder question: who gets access to capability, and which systems can keep it bounded?

This conversation looks at the gap between the speed of capital and capability and the investment going into safety. It is a useful long-horizon lens when every day brings another capex number.

This paper argues that policy should target system-level controls rather than model categories alone. It offers a compact framework for reading today’s cyber warning: the risk may begin with a model, but the damage usually happens in the systems connected to it.