In partnership with

> SPOTLIGHT

WHAT MATTERS TODAY

GitHub has added dedicated Agents secrets and variables for Copilot cloud agent, and it now lets teams share them at the organization level instead of configuring everything repo by repo. On the surface this looks like a small admin update, but the bigger signal is that coding agents are starting to get their own layer of permissioning, governance, and rollout policy.

Google says Gemini Robotics-ER 1.6 is now available through the Gemini API and AI Studio, with improvements in spatial reasoning, task planning, and embodied safety for robotics. This is not just another model launch. It is a sign that the same production-deployment logic now stretches beyond software agents and knowledge work into physical execution.

> SIGNAL HEADLINES

CAPTURE THE SHIFT

Anthropic says principle-based retraining removed a prior blackmail behavior in testing: Anthropic says its Teaching Claude why work pushed a previously reported blackmail behavior down to zero in the same testing scenarios after retraining the model to reason about causes rather than just optimize outcomes. This remains a company-led claim, but it is worth watching because it suggests controllability may be moving deeper into training methods rather than staying at the policy layer.

Anthropic raises Claude limits after signing new compute capacity with SpaceX: Anthropic says a new compute deal with SpaceX adds more than 300 megawatts of capacity this month and lets it raise limits for Claude Code and the Claude Opus API. The signal is straightforward: compute procurement is becoming part of the product experience, not just a backend detail users never see.

Airbnb says AI now writes 60% of its new code: TechCrunch reports that Airbnb says AI now writes 60% of the company’s new code, one of the clearest adoption datapoints yet from a large internet company. That matters because it pushes AI coding out of devtool hype and into day-to-day production reality at scale.

Google adds more inline source links and previews to AI Mode and AI Overviews: Google is giving AI Mode and AI Overviews more direct links inside the answer, plus previews that show users where a click will go. As AI becomes a default discovery surface, source transparency is starting to become part of product competition rather than just a nice UX detail.

Gemini API File Search becomes multimodal, filterable, and page-citable: Google has expanded Gemini File Search so it can search across multiple modalities, filter by metadata, and cite the exact page behind a result. That points to retrieval infrastructure moving from “find something relevant” toward “find the right thing and prove where it came from.”

> ONE PRACTICAL USE OF AI TODAY

Pin effort: xhigh to the right skill instead of raising reasoning across the whole workflow

The most useful practical takeaway today is not a new tool. It is the idea of forcing deeper reasoning only where it actually matters. If you use agents across multiple tasks, raising effort globally often makes everything slower and more expensive without improving the parts that need judgment most. Pinning effort: xhigh at the skill level solves the real problem: keep the default flow light, but force deeper thinking on the paths that justify it.

In 30 minutes, the goal is to separate the tasks that truly need reasoning depth from the ones that mainly need speed, then lock that decision into the skill instead of adjusting it by hand every time.

QUICK SETUP:

  1. Pick two or three agent tasks you run most often, such as research, review, draft, or debug.

  2. Decide which ones genuinely need deeper reasoning, usually the tasks with tradeoffs, synthesis, or evaluation.

  3. Open the matching skill file and add effort: xhigh only to those tasks.

  4. Keep the other skills at the default effort level so latency and cost do not rise across the whole workflow.

  5. Rerun the same sample task before and after the change, then compare output quality, speed, and consistency.

  6. If the output improves but the slowdown is too large, reduce the number of xhigh skills instead of raising effort everywhere.

MINI SCORECARD:

Skill

Main job

Real depth needed?

Better output?

Slower by how much?

Keep xhigh?

Research

synthesis / ranking

yes / no

low / medium / high

low / medium / high

keep / remove

Review

bug / risk audit

yes / no

low / medium / high

low / medium / high

keep / remove

Draft

rewrite / polish

yes / no

low / medium / high

low / medium / high

keep / remove

✦ HOW TO READ THE RESULT:

  • If the output only improves slightly but latency jumps hard, that skill does not deserve xhigh.

  • If research or review quality clearly improves, those are usually the best places to spend the extra reasoning budget.

  • If draft or formatting work slows down without getting meaningfully better, move those back to default quickly.

  • The goal is not to use a deeper model everywhere. The goal is to spend reasoning budget where judgment matters.

> PRESENTED BY BELAY

When Did Your Business Start Running You?

What started as ownership turned into obligation.

Now you’re in every meeting, decision, and channel… not because you want to be, but because things stall without you.

It’s not a capacity issue. It’s a structure issue.

The Freedom Framework shows you how to rebuild work flows, so you can step back without things breaking down.

BELAY U.S.-based Assistants help make that real by bringing ownership to execution, so your business doesn’t rely on you to function.

> WORTH READING

ANALYSIS & THESIS

Semafor frames both OpenAI and Anthropic as companies moving deeper into enterprise services, not just model access. Read it if you want to understand why the deployment layer is becoming a real battleground instead of a thin services wrapper around a model.

Why read: It sharpens the idea that control of implementation may matter as much as control of the model.

VentureBeat brings the agent story back to a concrete enterprise pain point: shadow adoption. Read it if you want to see why control planes, observability, and governance are quickly becoming mandatory once AI starts moving beyond the sandbox and into real work.

Why read: It gives the clearest enterprise framing for why agent governance is now part of the product stack.

a16z packs hyperscaler gains, slop surplus, call-center economics, AI app growth, and multi-vendor usage into one market frame. Read it because it offers a rare synthesis: AI is growing fast, but value capture, labor effects, and product behavior are still messy and uneven.

Why read: It gives a broader market lens for the same production story without pretending the economics are already settled.

Brookings offers a wider lens: practical AI advantage may not come from the model alone, but from the ability to pull more useful context into the right workflow. Read it if you want a concept that ties today’s signals together, because control planes, retrieval, workflow packaging, and enterprise deployment are all converging on the same question of who controls context best.

Why read: It connects the tactical product signals to a deeper capability thesis.

Together with 20,000+ builders and tech readers, cut through the noise & focus on what truly matters in AI.