Fort Worth 24

collapse
Home / Daily News Analysis / OpenAI is building AI agents for everything. Will everyone use them?

OpenAI is building AI agents for everything. Will everyone use them?

Sep 08, 2026  Twila Rosenbaum  8 views
OpenAI is building AI agents for everything. Will everyone use them?

How much control are you willing to give an AI model over your digital life?

Getting the most value from a large language model means giving it the keys to your inbox, your Slack account, your calendars, and your files. For a control freak or the AI-hesitant, that sounds like a lot. For Andrew Ambrosino, the lead engineer for OpenAI's desktop app, it is the only way to test the future. That is why the app now has access to, and control over, his inbox, Slack, phone, Notion, Figma, and more.

"If I'm asking it to write a document, is there a possibility that it's going to pull from a private DM on that subject and not know that it's not supposed to share some info? Yes," Ambrosino says. "I'll do it for the job. I will take the personal hit here and there if I have to. And I haven't had to."

Ambrosino works on OpenAI's biggest bet: ChatGPT Work, which was released last month and is available on the company's lowest subscription tier for $20 a month. The product is designed to allow white-collar workers to field AI agents, hooking LLMs up to the digital workflows used by accountants, investors, doctors, and everyone else whose day-to-day is dominated by a computer. OpenAI's marketing copy puts the goal succinctly: a world where intelligence goes beyond answering questions to helping everyone turn their biggest ideas into reality.

From coding to the corner office

For software developers, that shift is already happening, but it has been slow to spread to other departments. ChatGPT Work is a modified version of the company's Codex coding tool. It is meant to give non-engineers a version of the same functionality that software engineers already get from agents: an AI tool that does not just answer questions, but completes multistep projects on its own.

"In this new factor, ChatGPT can actually do entire, very complicated tasks for you all autonomously in a way that is delightful and safe," says Thibault Sottiaux, who leads OpenAI's core product work, including Work. "It's the very mission of OpenAI, to bring everyone along."

Commercially, that matters a lot. Agents that work for longer stretches burn through more tokens, making them more lucrative for OpenAI on a per-user basis. Reaching new professions is crucial, not just for OpenAI but for the industry at large. If coding has proven lucrative territory for AI labs, it is still a tiny subset of the professional work that AI tools need to enable if these companies are to justify their massive investment in training and computation. While labs have been focused on software engineers, vertical-specific competitors like Harvey for law and Clay for sales have been chasing those customers with a model-agnostic approach, meaning they will plug in whichever AI works best at the time.

Industry analysts see this as one of the major challenges facing OpenAI and its competitors. "If the labs cannot rapidly get ahold of the key complementary assets needed to scale AI in the market, value will accrue elsewhere," Christian Catalini wrote on Andreessen Horowitz's “It’s time to build” blog.

Making AI apps work for people who are not software engineers requires more hand-holding. OpenAI’s non-engineering workforce, like the communications and finance teams, started using Codex at a time when it was actively hostile to them. They were asked about code and shown empty diffs, technical readouts meant for software changes. "So, we started to make it more general purpose between February and now," Ambrosino says.

An OpenAI-backed study found that in June, 98% of OpenAI employees were using Codex, but just 17% of organizational subscribers and less than 1% of individual subscribers were using the agentic coding tool. That difference between near-total adoption inside the company and negligible adoption outside it is both the challenge and the opportunity for the company.

"The more value and the more utility that we generate for users, the more they will be willing to also pay for some part of that utility, and that’s how we’ve always seen ChatGPT as well," Sottiaux says. "You sit there and you’re like, ‘of course I want to pay $20 bucks a month for this,’ because the value that you get is so much more."

How to make AI intuitive

To understand the disconnect between internal enthusiasm and external adoption, it helps to understand what OpenAI’s engineers are building. Every LLM requires what engineers call a "harness" — the software wrapped around a model that decides what information it sees, which tools it can use, and how it presents its answers. If you want that model to do stuff — to become an agent — the harness gives it tools and instructions for using them on long-term tasks.

For developers, a command-line interface (CLI) that enabled LLMs to code was enough to change the way software was built and deployed. But most people are not using CLIs; there is a reason Windows replaced DOS. An agentic product that goes beyond software engineering is "going to be something that plays with the messy world of your life and your tools and websites that were built in 1995 and never updated," Ambrosino explains.

The team is trying to make functionality previously available only to coders as easy as prompting. Consider tools like Claude Code and Codex: they unleashed "vibe coding" by abstracting away all the actual software writing and letting users just tell the model what they want in a program. Now, OpenAI wants to make similar functionality as easy as a conversation.

"Without these products in front of the model, experts would know how to get the same results, but you wouldn’t get to a billion people using the thing," Ambrosino says. That trade-off between what power users need and what mainstream adoption requires plays out in internal debates at OpenAI, where some employees argue that a button is unnecessary if users can just ask the model directly. "We push back on that because it’s very early," Ambrosino says. "Discoverability matters in this phase, and at some point we won’t have the button."

Work has a few more buttons for selecting projects and plug-ins, but it aims for the same magic box interface as other OpenAI products. Ambrosino compares it to skeuomorphism, the fading practice of making digital tools look like the physical objects they replaced, such as a calculator app designed to look like a pocket calculator. "That stuff wasn’t just cringe design. That actually helped get people into this and make the transition."

OpenAI would not say how many people used Work versus Codex, but the joint app is used by just 20 million people, compared to more than a billion users the company says are prompting ChatGPT online.

Giving ChatGPT a license to skill

For now, OpenAI is pitching this tool as best suited for routine, data-intensive coordination tasks. Its employees are setting up weekly metrics reports and making spreadsheets into planning tools. There are venture capitalists using agents to assemble relevant communications and analysis about companies into investment memos, and operations teams spinning up bespoke dashboards and data visualizations. OpenAI’s CEO Sam Altman reportedly uses it to plan his vacations. One OpenAI engineer described asking the program to look at a Slack conversation about an engineering problem and "make some charts," and then receiving back a series of insightful plots.

"There is a deluge of information for the average worker or employee of any of these companies, including myself," says Akshay Nathan, who leads the product engineering team at OpenAI. "We’re actually quite limited by our ability to parse everything that’s available to us, and then take action on it. That information lives in all these system records tools like Salesforce. The value of ChatGPT is you already have access to this, but now you truly have access to it."

This could be the digital personal assistant that AI evangelists dream about. Like similar agentic products from rivals, ChatGPT Work links agents to your existing workspace — email, web browser, a slew of SaaS platforms — and puts that context to work for you.

When the system works, it can be impressive. One user asked ChatGPT Work to get a weirdly formatted preschool calendar from email and put it into Google Calendar, and it did, saving a lot of repetitive data entry. Another tasked it with financial analysis on publicly traded companies, and it delivered an auto-updating dashboard of metrics. It can create a queryable database of space launches, a task previously requiring Python scripts. It can even send a weekly email about new AI research posted at academic clearinghouses.

But trusting the model is not always easy. Setting up permissions for agents to access cloud drives can be confusing and circular. Many important settings are only available on the web app, forcing users to work across both desktop and mobile. Sometimes limitations are baffling: link ChatGPT Work to Google Calendar and it can create events, but not new calendars. Unless the effort level is set to high, you have got the worst intern you have ever worked with.

That is common advice from AI early adopters, who fear that frustrated newbies will give up. Joe Gershenson, an engineering lead for OpenAI’s harness, admitted that effort settings are not intuitive for new users yet. "There are things that we can do better to help them get the right level of reasoning," he says. Adding, "Watch this space."

Harder to measure than code

OpenAI faces another important challenge breaking into mainstream white-collar work: most workflows are not as measurable, or evaluable, as code. Software either works or it does not. A good presentation, business strategy, or sales pitch is not as easy to evaluate or trace. "One of the unique challenges with a product like this is just that it can really do anything," Ambrosino says.

OpenAI uses its benchmark GDPval, drawn from 44 occupations and hundreds of knowledge work tests, supplemented with user feedback. A less official answer is that product direction comes from OpenAI employees themselves. "We have to always parse out," Ambrosino says, "are we doing the workflow that everybody else will be doing, or are we weird?"

The early adopters of the app itself will create valuable traces with their actual usage. Much of the success of coding tools is built on similar data collection, assuming users do not opt out of making their data available for training.

The rivalry that drove OpenAI’s product design

Despite all the attention on the model interface, OpenAI’s engineers were reluctant to answer a fairly simple question: What sets Codex and ChatGPT Work apart from rival agentic harnesses intended for a mass user base?

"It’s going to be a really disappointing answer for you, and I’m sorry, but the honest answer is that I really don’t look at the harnesses that they’re building," Gershenson said. "The Mad Men ‘I don’t think about you at all’ meme comes to mind here."

Even if engineers claim not to look, product similarities suggest otherwise. The first thing ChatGPT Work asks new users to do is often to import data from a competitor. The rivalry is understandable. Anthropic’s Claude Code defined the market for AI coding and launched a revolution in how software engineers do their jobs. It is additionally frustrating because OpenAI had the idea first but did not quite harness it correctly.

When OpenAI first developed Codex as a web app, the engineers got over their skis, or "a bit more AGI-pilled," as Ambrosino puts it. They bet on the model being smart enough to handle a task entirely on its own, with minimal user input. Claude Code was oriented around a back-and-forth conversation with the user. If you gave it a problem, it would survey the possibilities and give you several options for proceeding. Once you chose, it would go a little further and then check back again, leaving less room for error.

Anthropic’s approach proved more effective, even if it demanded more work from users. "Our product was a little ahead of where the model and harness was at the time," Ambrosino says now. OpenAI eventually followed suit by adding more opportunities for users to interact with the model. That became the Codex known today, with desktop and mobile apps. Using download statistics as a proxy, Claude Code was more in demand until April of this year, but Codex has now taken a slight lead. Surveys of enterprise use also suggest OpenAI is catching up.

Part of that lead comes from getting product-market fit right. Part comes from complaints about safety restrictions on Anthropic’s models and compute shortages. OpenAI’s steps toward a more human-centric harness continue with ChatGPT Work, but engineers insist the key differentiator is the strength of OpenAI’s latest cost-effective models.

"The frustrating answer is that a lot of times it is the model," Ambrosino says.

What makes a good harness, anyway?

That explanation returns to the "bitter lesson" learned by AI researchers: a better general model is more important than specific domain experience. For true believers, the harness is a temporary crutch, not the moat. "You could get good results in the short term by adding a whole bunch of extras — ifs and thens and tools — but the next model is going to come out in a couple of months and make that obsolete," Gershenson says. His team focuses on the simplest ways to expose the model to the tools and context it needs, and no more.

"The goal of good harness engineering is to be more precise about what information the model really needs to solve your problem, because the models are getting better and better at doing that if you simply let them do their thing," Gershenson says.

There is an open question whether most people or models are ready for that. Some researchers who study AI tools in the workplace still see rivals as more user-friendly, noting that ChatGPT tends to want to do magic and just do it for you, while competing models do comparisons, show them, repeatedly ask for input and feedback, and run A/B tests.

Sottiaux and perhaps OpenAI at large disagree, arguing that the conversational nature of the app is better than learning how to use an application. "We definitely see that the world seems to be ready," he says. "This is why we’ve had incredible adoption."

Still, it is not clear that a model-specific harness is even the right bet for maximizing a model. Industry comparisons have shown that different harness and model combinations deliver different performance on coding benchmarks. One open-source harness, Pi, has been shown to outperform Codex while using the same underlying model. Pi has been used to build projects like OpenClaw and other cloud platforms.

Pi’s creator, Mario Zechner, says his intentionally minimalist harness is evidence that an AGI-pilled approach can work, at least for software engineers and coding tasks. What it lacks in explicit features, he says, is made up for by its ability to modify itself and build its own interfaces. He sympathizes with the challenge that OpenAI’s engineers face in expanding beyond engineers.

"Everything is coding agent shaped. The reason is that they only have training data for coding agent tasks," he says. "Say I’m in management, I make a decision today, and the outcome happens months later. You cannot capture that in a simple trace of a user and agent back and forth, so all of those kinds of tasks, anything that you don’t digitize, is inaccessible to a model to learn."

Like other open-source providers, he sees the big labs’ efforts to push their harnesses as a way to lock in users. "They need to own the entire stack; otherwise, they just become a model provider and then need to compete with Chinese models."

There is limited transparency into token spend and agent behavior in frontier labs’ harnesses. As with coding tools, uptake at the scale OpenAI hopes for will eventually force harder conversations about cost. For example, one user on a $20-a-month subscription used more than 80 million tokens in four days, which cost $65 according to the model’s own analysis. That is a subsidy of more than three times the subscription price for four days of casual use alone.

"We are working every day to push the frontier on efficiency," Sottiaux says, pointing to a recent 80% price cut for users of OpenAI’s Luna model. "If you wake up six months from now, you should be able to do all of the same tasks with less spend."

The other relevant question is whether these apps create a lock-in effect through data retention or the sheer pain of configuring access to all the plug-ins and their permissions. Inside OpenAI’s headquarters, the atmosphere was calm but slightly tense; these are people with a lot to do. The engineers were constantly monitoring their laptops and rushing between meetings.

Nathan, the head of the product engineering team, remains optimistic. The focus is on "the promise of the magic box, but I still think there’s too much complexity. I’m very optimistic that we can solve it, with the model and in a truly AI-native way."


Source: TechCrunch News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy