An empty chair and closed laptop beside a teal dot hovering over bounded task trays, with completed work awaiting human review.

Opening Note

Most of us use AI in a familiar rhythm: ask for help, inspect the result, then decide what comes next.

OpenAI’s new dots invite a different rhythm. Give an agent an ongoing responsibility, and it can keep working while you turn your attention elsewhere.

That sounds useful. It also raises a practical question: what would you trust an agent to keep doing without a fresh instruction each time?

This week, we’ll look at what OpenAI has introduced, what remains unproven, and how to test this kind of delegation on a manageable piece of work.

The Big Signal — What Would You Delegate to an Always-On Agent?

Imagine asking an AI assistant to investigate a bug. It examines the code, proposes a fix, and waits for your response.

Now imagine giving it a continuing responsibility: watch a particular feedback channel, identify recurring bugs in one component, and prepare proposed fixes for your review.

That second arrangement changes what you need to specify. Alongside the task, you need to explain what deserves attention, which actions are permitted, and when the agent should involve you.

On 29 September, OpenAI introduced dots: agents powered by GPT‑6 Astra, with their own cloud computer and access to connected applications. OpenAI describes them as able to pursue goals over time and learn from feedback. One illustrated use case is monitoring customer feedback and preparing tested fixes for review. These are capabilities described by the company; we have not independently tested them.

From a request to a responsibility

The potential benefit is continuity. If delegation works well, you spend less time reopening conversations, rebuilding context, and reminding the assistant what needs doing.

But “keep improving the application” is a poor first assignment. How should the agent choose between a cosmetic complaint and an intermittent payment failure? What happens when an apparent bug is actually intended behaviour?

A more useful trial would be narrow: monitor one component, gather supporting evidence, and prepare recommendations. Keep prioritization, merging and deployment with the team while you evaluate the results.

That gives you something concrete to assess: did the agent find worthwhile work, explain its reasoning, and reduce the effort needed to act?

What can happen while you are away?

OpenAI makes an important distinction. Unprompted background activity, which it calls proactive research, uses read-only tools. It cannot send messages, change connected-app content, or control a computer. Task execution is governed by permissions, built-in safeguards and configurable rules; an automatic review checks consequential actions against those controls.

So “always on” should not be read as unlimited permission to act.

For a team, the practical work is defining the assignment clearly:

  • Scope: which project, information and systems may it use?

  • Output: what should it bring back, and what evidence should accompany it?

  • Intervention: which uncertainty or action requires a person?

  • Evaluation: how will you count useful results, missed work and correction effort?

What remains unproven?

Dots are rolling out to Pro and Business Premium users in eligible markets. Enterprise access is an administrator-enabled beta, while specialist organizational dots begin through focused pilots.

A launch establishes availability and intended behaviour. It does not establish how consistently the agent will handle changing priorities, incomplete information or conflicting instructions over several weeks.

There is also a simpler alternative worth considering: a scheduled script or existing automation may handle a predictable responsibility adequately. An agent earns its place when interpreting changing information adds enough value to justify its cost and supervision.

Start with one bounded responsibility. Judge the quality of the work it brings back—and how much attention that work still needs.

Worth Knowing

Codex: the development environment becomes part of delegation

At DevDay on 29 September, OpenAI announced expanded cloud execution for Codex, with reusable development environments and access from different devices. Those environments can provide a shared setup with approved settings and permissions.

That matters because an agent needs more than repository access. It also needs the right dependencies, commands and permissions to build and test a change. A consistent environment could reduce repeated setup and make results easier to reproduce.

For a team trying this, check whether a fresh cloud task can run the same acceptance checks as a developer’s machine—and whether its proposed changes arrive with enough evidence to review.

MCP events: connected applications can initiate work

OpenAI also announced support for the proposed MCP Events specification, allowing events in connected applications to start automations. For example, a new task on a project board could trigger preparation of a draft plan.

MCP—the Model Context Protocol—provides a common interface between AI applications and external tools or data. Its Triggers and Events working group is developing mechanisms for servers to notify clients proactively, including delivery and ordering behaviour. A product implementation should not be mistaken for a finalized, universally supported standard.

The practical implication: an automation may need to handle duplicate events, changed priorities and retries. “Something happened” starts the workflow; it does not settle what the workflow should do.

A model retirement belongs on your maintenance calendar

A study published on 25 September examined model-migration commits in open-source applications. It estimates that 82% of migrations associated with retired models were committed after the shutdown date.

That is a warning about dependency maintenance, with an important limit: a late repository commit does not independently establish when—or whether—a deployed application stopped working.

If your application depends on a hosted model, record the model identifier, monitor retirement notices, and assign someone to test its replacement before the deadline. Changing a model name may be a small code edit; checking whether the replacement preserves useful behaviour can take considerably more work.

From the Engineering Desk — A Completed Run Is Only the Beginning

This week, I’ve been testing a system that turns an information request into a research plan, gathers evidence, and prepares findings.

Getting a live run to complete was encouraging. Then I looked inside the result: what request had the system understood, what had it searched for, and which evidence supported its findings?

One question stood out: why had it produced so many findings from the same source?

That is not automatically a problem. If someone asks about a regulator’s new rules, one authoritative document may contain several relevant facts. Breaking those facts into separate findings can make the answer easier to inspect.

But several findings from one document are still grounded in one document. They do not become independent corroboration simply because the system presents them separately. For a question that needs competing perspectives or confirmation across sources, that concentration deserves attention.

It reminded me to separate three checks:

  • Execution: did the workflow finish?

  • Coverage: did it gather evidence for the important parts of the request?

  • Support: are the findings relevant, traceable and sufficiently supported?

A completed run answers the first question. It does not settle the other two.

My next step is to try more varied requests and inspect where the research is strong, narrow or incomplete. The useful milestone is a result worth relying on—and enough visibility to understand its limits.

Worth Your Time — Find Out Where the Agent Failed

If you are evaluating an agent that gathers information and uses tools, take a look at AgentHop, published on 28 September.

The benchmark examines performance across four dimensions: retrieval, synthesis, tool use and resource management. That distinction is useful: an incorrect answer might come from missing evidence, misreading good evidence, a failed tool call or an exhausted budget. Each calls for a different fix.

Its controlled scientific-question setting limits how far the results generalize. The transferable idea is the diagnostic approach: measure enough of the workflow to know what needs improving.

Before You Go

Think of one responsibility you repeatedly return to: checking feedback, investigating recurring failures, or keeping a technical brief current.

What would an agent need to know to handle it usefully? Consider the scope, the evidence you would expect, and the point at which it should involve you.

What is one ongoing responsibility you would like to delegate to an AI agent—and what would make you trust its work?

Reply and tell me. I’d like to hear where this would help in your working week.