Only the Best artificial intelligence agents An experiment which challenges the idea that AI will replace office workers en masses has shown that most freelancers are not very good at their online work.
The Remote Labour Index is a benchmark created by Scale AI researchers and CAIS (Center for AI Safety), an independent nonprofit. It measures how well frontier AI models can automate work that has economic value.
Researchers tested several AI agents and discovered that they could do less than 3% of work. The best AI agents were only able to earn $1.810 from a potential $143,991 earning. The researchers examined several tools to determine which was most effective. They found Manus by a Chinese startup with the same title, Grok by xAI (also known as xAI), Claude by Anthropic, ChatGPT and OpenAI from OpenAI were the top four.
“I should hope this gives much more accurate impressions as to what’s going on with AI capabilities,” Dan Hendrycks says, “Director of CAIS. He also says that although some agents’ performance has improved over the last year, it does not necessarily mean the improvement will continue.
AI’s spectacular advances have fueled speculation that AI will soon replace workers in large numbers. Dario Amodei CEO of Anthropic suggested in March that 90% of the coding work is automated. would be automated In a matter months.
AI waves in previous years have led to faulty predictions, such as about the displacement of jobs. imminent replacement of radiologists AI algorithms are a great way to improve your business.
Researchers generated freelance tasks by using verified Upworkers. These tasks include graphic design, video-editing, game development and administrative duties like data scraping. Each job was described, along with the necessary files and a sample of what a person would have produced.
Hendrycks claims that AI models are getting better, but they have not improved. at codingMath is a great way to learn. logical reasoning In the last few years, many people still have difficulty using different tools or performing complex tasks that include multiple steps. “They don’t have long-term memory storage and can’t do continual learning from experiences. They can’t pick up skills on the job like humans,” He says.
OpenAI, in its benchmarking of economic research published by OpenAI last September called GDPvalThis is a measure of work that has economic value. GDPval claims that frontier AI models like GPT-5, which are based on the latest AI technology, can perform 220 office tasks with human-like accuracy. OpenAI declined to comment.

