OpenAI Dots Review: The Always-On Agent That Can't Retrieve Its Own 2FA Code

Sources

Yesterday at 17:07 UTC OpenAI shipped dots: always-on agents powered by GPT-6 Astra, each with its own cloud computer and browser, reachable through ChatGPT, Slack and Teams, working toward your goals 24/7. The launch thread hit 652 points and 500+ comments in a day — and unlike most agent launches, the criticism inside that thread is not vaporware skepticism. It is hands-on users describing exactly where the product's autonomy ceiling sits, a security researcher documenting a TLS interception proxy inside the dot's cloud browser, and Pro subscribers reading an email that halved their existing usage allowance the same week dots arrived.

This review is built from OpenAI's primary documentation set — the dots docs, the workspace admin guide, the published auto-review evaluation and the open-source guardian policy that backs it — plus the launch-day HN evidence. The short version: dots is the most honestly-engineered personal agent a major lab has shipped, because almost nothing in it is actually autonomous. Whether that is a feature or a disappointment depends entirely on which of the two products you thought you were buying.

Executive Scorecard

DimensionScoreVerdict
Reliability 6/10 Persistent cloud computer with deliberate pause/wake cycles, parallel background agents, and durable notes across channels. But confirmation loops repeat, 2FA walls kill end-to-end automation, and local tasks die silently whenever the connected laptop sleeps.
DX 7/10 Setup on desktop only, no mobile web, one personal computer at a time. Context genuinely carries across ChatGPT, Slack and Teams. The custom-rules rule-checker rejects broad authorizations, so the autonomy you want to configure is often the autonomy it refuses to grant.
Cost 4/10 "Your first dot is included" — but the same email cut Pro 200 allowances by half (20x to 10x Plus, 200 to 100 weekly GPT-6 Pro messages) from October 30. The deeper-work allowance past month one is undisclosed. Pro 500 exists at $500/month. The precedent is the warning.
Security 6/10 Four enforcement layers (app permissions, custom rules, auto-review, run monitoring with a kill switch) and a genuinely public reviewer policy. But the cloud browser's egress runs through an OpenAI TLS interception proxy, disconnecting an app does not delete what the dot already ingested, and enterprise model controls do not apply to dots at all.

What Actually Shipped

Dots are not a chat mode. Each dot is a standing GPT-6 Astra agent instance with three durable assets: a cloud computer (files, software, and a browser whose sessions persist between periods of use), a notes store that survives channel switches, and a set of connections — ChatGPT plugins (the announcement claims 4,000+ reachable apps), a Slack or Teams presence, and optionally one of your laptops. It can spawn background agents that run in parallel, open separate cloud threads, and delegate to Work or Codex tasks — local ones on your connected computer, cloud ones in a Codex environment you pre-provisioned.

Availability, straight from the docs: Pro users over 18 outside the EEA, UK and Switzerland; Business Premium worldwide; Enterprise worldwide but off by default until a workspace admin enables it. Texting is a limited US-only Pro beta through a third-party provider. Teams is an invite-only alpha. Specialist dots — the enterprise play, with their own identity, credentials and systems-of-record integrations, plus a Microsoft Agent 365 governance integration — are a preview with "focused enterprise pilots" only.

The pricing story is where launch week turned sour. A Tell HN from a Pro 200 subscriber reproduces the email: starting October 30, 2026, Pro 200's included usage in ChatGPT Work and Codex drops from 20x to 10x the Plus allowance, and the weekly GPT-6 Pro chat limit falls from 200 to 100 messages — price unchanged, sweetened with a one-time 62,500-credit grant ($2,500 value, expiring December 31, 2026). The Pro tiers page confirms the new ladder:

PlanPriceDotsAstra UltrafastIncluded usage vs. Plus
Pro 100$100/moYes (non-EEA/UK/CH)NoTier allowance, per Pro tiers page
Pro 200$200/moYes (non-EEA/UK/CH)No10x from Oct 30 (was 20x)
Pro 500$500/moYes (non-EEA/UK/CH)Yes (8x token generation)25x
Business Premiumper-seatYes, worldwideNoPer-plan allowance

Two fine prints matter more than the numbers. First: conversations with your dot and work it does directly "don't count toward your ChatGPT usage allowance" — but any task it starts or manages in Work or Codex counts toward those products' limits as usual. Second: your plan "includes an allowance for deeper work, with extended limits for the first month after launch." After that month, the allowance is undefined. A product whose headline feature is 24/7 operation ships with an undisclosed consumption cap on that operation.

Architecture Mechanics

The dot is an orchestrator, not a single loop. From the docs, the steady-state topology looks like this:

        you (ChatGPT / Slack / Teams / voice)
                       |
             [ dot orchestrator ]  <-- persistent notes
               |       |       |        (survive channel switches,
               |       |       |         separate from ChatGPT memory)
   background  |       |       |
   agents xN    |       |       |
   (parallel,   |       |       |
   read-only     |       |       |
   proactive     |       |       |
   research)     |       |       |
                 |       |       |
        [cloud threads] |  [local Work/Codex tasks]
         (visible in    |   (laptop online + ChatGPT app open,
          desktop/web/  |    only ONE computer, connection
          mobile apps)  |    persists between tasks)
                        |
              [ cloud computer ]
              - own files, software
              - own browser sessions
              - egress via OpenAI TLS
                interception proxy  --> monitoring system
                                        (can pause/stop the dot)
                        |
              [ plugins: Gmail, Drive, GitHub, ... ]
                 each with own account
                 + permissions you set
                        |
        every account-affecting action
                        |
              [ action review ]
        instructions + custom rules + safety
        -> proceed / ask you / hand off to you

Loop mechanics worth knowing as an operator:

Initiative is read-only by design. When you are not talking to it, the dot does "proactive research": background agents read your connected apps, form private notes, and surface suggestions. The docs repeat three times that these research tools cannot send messages, change app content, or control a browser or computer. Everything actionable still routes through the approval stack. "Always-on" means always-reading plus whatever you have explicitly authorized — it does not mean always-acting.

State is split four ways. Conversation context (per-interaction selection), ChatGPT's saved memory (shared, governed by your memory settings), the dot's own notes (separate; changing ChatGPT memory settings "doesn't necessarily change the notes your dot has already made"), and website sessions held in the cloud browser (persist until sign-out or site expiry). A task the dot delegates gets "instructions and context from your dot for that work" — it does not inherit every conversation your dot has had. That boundary is good hygiene, and it also means delegated tasks can be missing context you assumed they had.

Scheduling is event-driven by default, cron-like on request. The dot "can decide when to pause and wake up" between conversations — no fixed cadence unless you ask for one. If timing matters, you specify a schedule (with time zone and end date) and it lands in the Scheduled view. Where a connected service supports it, you can trigger on events — "investigate new bug reports in a connected Slack channel" — but adding the dot to a channel alone establishes nothing; it confirms the monitoring task when asked.

Local delegation is the fragile edge. The task-type matrix from the docs, with its operational preconditions, is the part platform engineers should tape to the wall:

Task typeWhere it runsHard precondition
New local Work or Codex taskYour connected computerComputer online, ChatGPT app open, access enabled
Existing local Codex taskYour connected computerSame — identify task and changes
New cloud coding taskCodex cloud environmentEnvironment pre-created in Codex with repo + setup
Dot-created taskOriginal computer or cloud envFollow-ups go to the task's original environment

Only one personal computer can be connected at a time, the connection persists between tasks, an "Offline" status means unavailable rather than revoked — and switching the dot's selected computer does not migrate existing tasks. If the laptop sleeps, local work stalls; if the website blocks cloud browsers (some do), the fallback is a separate local task, which does not carry the cloud session. Every one of these is a support ticket waiting for a Monday morning.

The Enforcement Stack: Custom Rules and Auto-Review

Dots' safety architecture is the most substantive thing about the launch, because OpenAI published its internals. Four layers stack:

Layer 1 — plugin permissions. Your existing ChatGPT app controls apply; you can grant read email without send email. Connecting a channel grants nothing else.

Layer 2 — custom rules. Settings → Personalization → Custom rules. Each rule names an action and picks one of four behaviors:

Rule behaviorMeaning
Take action without askingThe dot proceeds with no approval prompt
Take action when you say soProceed only when you explicitly request it in the prompt; otherwise ask first
Ask before taking actionApproval required before every occurrence
Hand off to youThe dot never does it; you do

The docs are explicit that rules are "instructions your dot tries to follow" — they do not override built-in safety requirements, do not grant app access, and do not remove mandatory confirmations. Admins can disable custom rules workspace-wide; when disabled, saved rules stop applying entirely.

Layer 3 — auto-review. Before any action that "could affect your accounts or share information," a separate reviewer agent — not the dot — checks it against your instructions, permissions, custom rules and safety requirements, then decides: proceed, ask you, or hand off. The announcement links this directly to the auto-review system shipped in Codex in April, and that system's reviewer policy is public in the openai/codex repository. It is worth reading, because it is a real risk taxonomy, not marketing: data exfiltration (deny sensitive payloads to untrusted destinations, tracing derived data back to its source), credential probing (reading credentials from unintended sources like browser profiles is high risk; normal auth-service flows are routine use), persistent security weakening (prod or wildcard control disablement is high; narrow, task-bounded disables are medium), and destructive actions (broad refspecs, branch deletion, protected-branch pushes are high; bounded local edits are low).

Two design details stand out. A denial is not a dead end: the reviewer returns a rationale and the main agent is instructed to find a materially safer path or stop and ask. And there is a rejection circuit breaker — after 3 consecutive denials, or 10 denials in a rolling window of 50 reviews in one turn, the turn is aborted with an interrupt rather than letting the agent grind against the wall. The published evaluation (reviewer: GPT-5.4 Thinking, low reasoning) reports a 99.1% auto-approval rate on escalated actions, 90.3% recall on overeagerness cases, 99.3% on prompt-injection exfiltration attempts, and 96.1% on MonitoringBench hard — with the honest caveat that red-teaming found cases where the reviewer could be misled.

Layer 4 — run monitoring with a kill switch. A monitoring system watches for malicious instructions and unsafe behavior and "can pause or stop the dot's work." What that monitoring is physically able to see is the next section.

Now the part no review should let you skip: does this stack actually let you delegate? Launch-day hands-on reports say the ceiling is lower than the demo reel implied. One Pro user's two-hour test described a confirmation loop — the dot filled a booking form on its cloud computer and then came back to ask whether it should book, after the user had already confirmed twice. The same user found that when a login triggers a 2FA code to email, the dot refuses to fetch and enter the code even with mailbox access granted, no matter how explicitly instructed — the verification stays with the human, full stop. And the custom-rules UI includes a rule-checker that evaluates your proposed rule for scope: this user's "proceed on routine, reversible, low-risk actions" rule was rejected as overly broad, and a Slack auto-reply rule was rejected outright. Two rules accepted out of many attempts.

Read charitably, that is a launch-period caution posture working exactly as designed — every rejection above is a category where an agent has burned a user before (payments, channel impersonation, credential reuse). Read literally, it means dots currently cannot be configured into the autonomy its marketing shows, and anyone evaluating it for real delegation should test their specific approval loop before assuming the "24/7" includes actions rather than research.

The Interception Proxy in the Cloud Browser

The most important security fact about dots is not in the announcement. It is a launch-day test by HN user chen_dev, who ran a TLS handshake from inside the dot's cloud environment — via curl, the bundled browser, and a freshly downloaded Firefox — and got the same answer every time:

$ site under test: https://gmail.com
  Subject: CN=gmail.com
  Issuer:  O=OpenAI, LLC; CN=openai.com
  Certificate verification: passed (inside the cloud browser)
  Response: 301 redirect to mail.google.com

OpenAI-operated TLS interception on the dot's browser egress. Every HTTPS session the cloud browser holds — including sessions you signed into through the "private form" — is readable by OpenAI's proxy, by construction. That is presumably how the monitoring system does its job: it cannot pause a dot for dangerous behavior if it cannot see the behavior, and agent-visible logs alone are not trustworthy when the agent may have been hijacked. As a control, it is defensible; centralized visibility into agent egress is something every serious deployment builds. As a privacy fact, it should be stated the way I just did: the "saved passwords without exposing them to the model" guarantee is about the model. OpenAI's interception path sees the credentials, the sessions, and the traffic. If your compliance posture says a third party must not be able to read sessions into your systems of record, dots' cloud browser is disqualified from holding those sessions — use plugin scopes, or connect a local computer where policy allows, and keep the cloud browser on public, low-stakes work.

The rest of the data lifecycle has the same shape: reasonable defaults, sharp edges on deletion. Disconnecting an app does not delete information the dot already ingested — the only purge path for its memories is deleting the dot. Pausing stops the current main task only; delegated tasks keep running and schedules keep firing until you stop each one separately. Stopping work never undoes completed actions; deleting the dot does not recall messages or revert app changes. The admin guide's offboarding guidance is effectively "review saved memories and reset before you deprovision the human" — correct, and a checklist your HR-driven identity lifecycle does not have today.

Two more enterprise facts from the admin guide deserve plain language. Model controls do not apply to dots — your workspace's model allowlist and defaults govern nothing here; GPT-6 Astra is the engine, and access is governed only by the dots capability toggles. And if any cloud policy has enforce_residency enabled, local computer access for dots is unavailable outright — a data-residency tripwire that will surprise the first multinational that tries it. Admins get real RBAC (Use dots, Add dots to Slack, Allow local computer access, Use custom rules, Cloud browser/network/computer use, Use password manager), a Compliance API for investigating dot activity whose coverage you are told to confirm before relying on audits, and an Analytics API for adoption metrics.

Critical Failure Modes

Confirmation-loop divergence. When review, custom rules, and site friction disagree, the dot can orbit an action: request → confirm → execute partially → request again. The auto-review circuit breaker caps this inside a single turn (3 consecutive denials abort), but a loop of approvals — like the booking-form repeat — has no breaker. Your only control is watching Activity and redirecting.

2FA and verification walls. Any flow that touches a verification code, a password change, or an equivalent "sensitive task" terminates at the human. With mailbox access the dot still will not fetch the code. This is correct security and a hard ceiling on end-to-end automation for exactly the flows (bookings, payments, admin consoles) people buy always-on agents to handle.

Autonomy is capped by a rule-checker you cannot argue with. The rules UI evaluates proposed rules for scope and rejects broad authorizations with a reason. You cannot ship "act on routine reversible actions" — the category a power user actually wants. Until that posture loosens, dots is a research-and-draft agent with gated execution, and evals should be run against the rules UI, not the demo video.

Disconnect does not mean delete. Ingested app content and the dot's derived memories survive app disconnection; the only purge is dot deletion, which also destroys its task history and notes — a blunt instrument that makes selective retention impossible.

The local-computer dependency. One computer, must be online, app open, access enabled. Laptops sleep. Expect silent task stalls, and expect the "Offline" status to be interpreted as "disconnected" by users who did not read the difference.

Usage-allowance drift. Dot conversations are unmetered, but month one's "extended limits" expire into an undisclosed allowance, and delegated Work/Codex tasks draw on product limits that OpenAI just demonstrated it will rebalance mid-subscription. Any workflow designed around a dot's continuous capacity is designed around a number OpenAI has not published — and has shown willingness to cut.

Monitoring asymmetry. The interception proxy means OpenAI's visibility into the dot's browsing exceeds yours. You see Activity; OpenAI sees the traffic. For regulated work, assume the cloud browser is OpenAI-monitored infrastructure and route accordingly.

The Ecosystem Context: Same-Day Reversals and Fast Clones

Two ecosystem signals landed within 18 hours of dots, and both matter more than the usual launch-week thinkpieces. First, Earendil Engineering published "You Said No MCP!" — pi.dev, which had proudly refused MCP support, brought it into the core the same day dots launched. Their reasoning is a genuine engineering position, not a capitulation: modern MCP works better when tools return structured data and compose through a code sandbox ("codemode") instead of dumping text into the context window. The agent world is converging on structured tool composition as the pattern — dots' plugin stack, pi's codemode, and the MCP-in-production patterns we covered earlier are all answering the same context-economics question.

Second, the clone count. Within a day, GitHub hosted at least three open-source "OpenAI Dots alternative" projects (Anil-matcha/open-dots, diggerhq/opendots, ychampion/melete) — and the serious self-hosted alternative already exists and is not a clone: OpenClaw, the agent framework running root on your own hardware. The honest trade: OpenClaw gives you real autonomy and real blast radius — its history includes a privilege-escalation CVE and providers restricting accounts that drive subscriptions through it — while dots gives you a fenced, monitored, reversible-by-default experience whose fences are the product. Teams choosing between them are choosing who they trust more with a credential: themselves, or OpenAI's proxy.

For readers tracking the broader agent-tooling landscape, we've covered adjacent ground in our AI coding agents review, AI code review agents, and AGENTS.md adoption.

Who Should Skip This

EU, UK and Swiss teams — you have no decision to make yet; the rollout excludes you, and the training-default posture of proactive research is plausibly why.

Anyone with data-residency enforcement — enforce_residency blocks local computer access, and the cloud computer's interception proxy puts all its browsing through OpenAI. If your systems of record must not be readable by a third party, the cloud browser is out of scope, full stop.

Teams that need guaranteed audit trails. The Compliance API's coverage of dot activity is explicitly something you must "confirm before relying on it for an audit." Until you have confirmed coverage for dot-created threads, delegated tasks and background agents, dots is not your regulated-workflow agent.

Buyers of the "handles everything" pitch. If your use case is outbound action — sending, booking, paying, publishing — at volume, launch-day evidence says the approval stack stops you at exactly those actions. Buy dots for research, monitoring, drafting and code-prep with gated execution; do not buy it as an autonomous actor.

Self-hosters. If the interception proxy, the undisclosed allowance, or the Pro 200 cut cross a line for you, the alternative is not a worse product — it is a different trust model. OpenClaw on your own metal, or a harness you wire yourself, remains the path where the operator holds root.

Verdict

Dots is a well-engineered answer to the right question — "how do you let an agent run continuously without giving it the keys" — and the publication of the guardian policy and its evals is genuinely above-and-beyond industry practice. The four-layer enforcement stack, the read-only initiative model, and the kill switch are the correct skeleton for always-on agents, and the specialist-dots preview (own identity, own credentials, admin-managed responsibilities) is the first credible sketch of agent IAM — service accounts for AI — from a major lab.

But the honest review is that dots ships as a research agent with gated hands, not the "handle everything" agent from the keynote. Its initiative is read-only; its execution is approval-gated in exactly the categories the marketing emphasizes; its rule system refuses the broad authorizations that would change that; its most impressive autonomy demos terminate at a 2FA wall by design. The pricing context — Pro 200 allowances halved in the same email, an undisclosed post-month-one allowance, a $500 tier for the full-speed version — tells you the meter is coming, even if the marketing says the conversations are free.

For platform engineering teams, the right eval is small and specific: one dot, non-production apps, one recurring monitoring task against a Slack channel or feedback source, one delegated Codex task in a cloud environment — then read Activity for a week and try to write the custom rules you actually need. If the rule-checker accepts them, you have found a real tool for ambient engineering support. If it rejects them, you have learned the actual product boundary before it learned your calendar. Either way, keep the cloud browser off your systems of record, and treat everything it browses as readable by OpenAI — because as of this launch, it is.

References & Further Reading