Say you ask an AI to send this week's expense report to your manager. A chatbot writes you a polite email and stops there. An agent is supposed to open the spreadsheet, check the totals, attach the file and press send. The second job sounds like a small step up. In practice, it's where most of these systems still trip.
Tech companies have spent the past couple of years selling that second version, so it's fair to ask how close it really is.
AI agents are built to carry out tasks across apps and websites, not only reply in a chat window.
What Makes an Agent Different
A chatbot produces text and waits for you. An agent takes a goal and works toward it on its own. It can use tools, click through websites, run code and check whether its last step worked before moving to the next. The International AI Safety Report describes such systems as ones that act independently and interact with their environment to reach a goal.
Where They Have Clearly Improved
The speed of progress is real. Researchers often measure agents by how long a task would take a human, then ask what length an AI finishes correctly half the time. According to the Safety Report's update, that length grew from about 18 minutes to over two hours in one year. Most of this testing is on software engineering and reasoning tasks. Coding is the strongest area for a simple reason: code either runs or it doesn't, so mistakes are easy to catch. Keep in mind that a 50 percent score means the agent still fails the other half of the time. The same report says agents are used in limited ways today, for web search, software development and trip planning, and that results vary from one area to another.
Where They Still Struggle
Ordinary office work is harder. In a benchmark called TheAgentCompany, researchers built a pretend software firm with 175 tasks, such as analysing spreadsheets, arranging meetings and messaging simulated colleagues. The best agent finished about 30 percent of them on its own. The Safety Report adds two more data points. In realistic customer-service simulations, the best agents completed fewer than 40 percent of tasks. On open-ended web jobs like planning a trip or making a purchase, the best model succeeded 12 percent of the time.
A newer paper, Workspace-Bench, points to a reason. Agents cope with clicking through screens, but they struggle when a job depends on many scattered files that have to be connected.
Why One Slip Ruins a Long Job
Jobs are made of many steps, and errors add up. If an agent gets each step right 95 percent of the time, twenty steps in a row succeed only about 36 percent of the time. That's plain arithmetic, not a measured result, but it helps explain why short tasks work and long ones fall apart.
The Safety Report also notes that agents can't build up knowledge of a workplace over time the way a human colleague does. Cost matters too. The Workspace-Bench paper quotes a survey in which 49 percent of enterprises named the cost of running agents as their biggest obstacle to scaling them.
What's Proven and What's Marketing
Be careful with the numbers above. The 30 percent and 12 percent scores come from tests on older systems, and newer ones probably do better. Benchmarks also use simulated companies, which are tidier than real ones. Phrases like "digital employee" come mostly from vendors, and the research cited here shows no case of agents replacing a whole role.
What the evidence supports is narrower. Agents can take on parts of a job: short, clearly defined tasks where someone can check the result. Treat one like a fast intern. Give it clear instructions, review what it hands back, and widen its responsibilities only when it has earned them.