DenaurumDENAURUM

OpenAI's Inbox Gamble: Why AI Agents Still Can't Replace Your Workflow

OpenAI handed an AI full access to an engineer's email, Slack, and phone to test whether agents can actually run your work life—but adoption outside the company sits below 1%.

Key takeaways

  • Andrew Ambrosino, lead engineer for OpenAI's desktop app, admits the AI could leak private DMs without realizing it's sharing sensitive information.
  • Inside OpenAI, 98% of employees use Codex; outside the company, only 17% of organizational subscribers and less than 1% of individual subscribers have adopted it.
  • OpenAI's ChatGPT Work app reaches 20 million users combined with Codex, compared to over 1 billion people using ChatGPT online—a gap the company frames as both challenge and opportunity.
  • The bottleneck isn't model intelligence but interface design and discoverability; Ambrosino argues buttons and UI guidance matter until users naturally know what to ask AI agents to do.

OpenAI just handed an AI the keys to someone's entire digital life: email, Slack, phone, Notion, Figma—the works. No breach. No accident. Voluntary experiment.

Andrew Ambrosino, the lead engineer for OpenAI's desktop app, is the guinea pig. And he's doing it to answer a deceptively simple question: Can AI agents actually run your work life, or is that promise still miles ahead of reality?

The Problem No One's Talking About

Ambrosino told TechCrunch something most AI companies won't say out loud: yes, the AI could leak your private DMs without realizing it's doing something wrong.

When asked about the risk of his AI pulling sensitive information from a private conversation and sharing it in a document, Ambrosino's answer was blunt: "There is a possibility that it's going to pull from a private DM on that subject and not know that it's not supposed to share some info." His response? "I'll do it for the job. I will take the personal hit here and there if I have to."

That's the trade-off at the center of OpenAI's biggest new bet: ChatGPT Work, released last month at $20 a month. It's built on Codex, the company's coding tool, repurposed for people who've never written code in their lives. The goal sounds grand—intelligence that goes "beyond answering questions" and turns "biggest ideas into reality." But the execution reveals how messy the gap between promise and practice still is.

The DOS-to-Windows Problem

Software engineers have lived this future already. Give them a command-line agent, and it transforms how they build. But most people don't use command lines. There's a reason Windows replaced DOS.

That's OpenAI's real challenge: not building a smart model, but building something that survives what Ambrosino calls "the messy world of your life and your tools and websites that were built in 1995 and never updated."

When OpenAI's non-engineering staff—comms, finance—first used Codex, the tool treated them like developers. It threw error messages at people who don't code. Ambrosino says it was "actively hostile" to them. Between February and now, the team had to rebuild it entirely to be general-purpose. Not smarter. Just less confusing.

Follow the Token Burn

Here's what explains OpenAI's entire strategy: longer, more complex agentic tasks burn through more tokens. More tokens per user equals more revenue.

Pushing agents into accounting, law, sales, medicine—into every white-collar job—isn't just altruism. It's the business model. Coding, as lucrative as it's been for AI labs, is still a small slice of professional work. If OpenAI is going to justify the money it's poured into training and computing, it needs White-collar America broadly, not just software engineers, using this daily.

Industry analyst Christian Catalini nailed it: "If the labs cannot rapidly get ahold of the key complementary assets needed to scale AI in the market, value will accrue elsewhere." Translation: build the tool people actually use in their workflow, or someone else will.

The Number That Matters

Inside OpenAI? 98% of employees use Codex.

Outside OpenAI? Just 17% of organizational subscribers use it. For individual subscribers paying for ChatGPT? Less than 1%.

That gap—near-total internal adoption versus negligible external use—is, in OpenAI's own words, both "the challenge and opportunity." ChatGPT Work and Codex combined reach 20 million users. Compare that to over 1 billion people using ChatGPT online. That's not a rounding error. That's a chasm.

It's Not About Intelligence. It's About Buttons.

The twist? The bottleneck isn't model capability. It's interface.

Ambrosino compares what his team is building to skeuomorphism—those calculator apps designed to look like physical pocket calculators even though they didn't need to. "That actually helped get people into this and make the transition," he says. The model might already be capable enough. Most people just don't know how to ask, don't trust it, or can't find the button that shows what's possible.

There's even internal debate at OpenAI about this. Some employees argue buttons are pointless—if the model's smart enough, just ask it. Ambrosino pushes back: "Discoverability matters in this phase, and at some point we won't have the button. But not yet."

So the honest answer to "can AI agents do everything?" right now is: technically, maybe. But almost nobody knows how to ask.