ON
← Back to feed
Has a Chinese physical AI start-up manipulated a global ranking to beat Nvidia?
HK🏛️ PoliticsCenter13 hr. ago

Has a Chinese physical AI start-up manipulated a global ranking to beat Nvidia?

A Chinese AI startup, Spirit AI, briefly claimed the top spot in a global robotics benchmark called RoboArena, surpassing Nvidia. The achievement sparked controversy due to concerns about benchmark integrity. Shortly after, the benchmark's creators revised their evaluation methods and removed Spirit AI and several other models from the rankings, citing 'benchmark hacking.' The incident highlights ongoing tensions between U.S. and Chinese tech firms in AI development and raises questions about the reliability of such benchmarks. The RoboArena, developed with Nvidia and academic institutions, assesses how well robot software translates digital commands into physical actions.

A Chinese physical AI start-up has sparked controversy after briefly surpassing Nvidia in a global benchmark for robotic intelligence, only for its achievement to be rescinded following accusations of manipulating the evaluation process. The incident highlights the growing tensions between U.S. and Chinese firms in the race to advance artificial intelligence technologies, particularly in the field of embodied intelligence, where robots must perform tasks in unpredictable environments. The dispute began in June when Spirit AI, a startup based in Hangzhou, Zhejiang province, unveiled its latest model, Spirit v1.6, at the start of the month. The company claimed the model achieved the highest score on RoboArena, a widely recognized benchmark for assessing how well AI-driven robots can execute complex physical tasks. This result temporarily placed Spirit AI ahead of Nvidia, a leading American technology corporation known for its expertise in AI hardware and software. The benchmark, which Spirit AI described as “the Olympics of embodied intelligence in North America,” drew considerable attention due to its high-profile participants and the prestige associated with winning. However, the success was short-lived. Within days, the RoboArena team announced a major revision to its evaluation criteria and removed multiple models, including Spirit v1.6, from the official rankings. According to statements made by Pranav Atreya, a key contributor to the project and a doctoral candidate at the University of California, Berkeley, the decision followed an investigation into potential violations of the benchmark's rules. Atreya noted that the team had “retroactively removed evaluations from organizations who [it] found to be engaging in benchmark manipulation.” While he did not specify which entities were affected, the removal of Spirit AI’s model suggested that the startup had been implicated in some form of unethical behavior. The controversy surrounding RoboArena intensified further when it was revealed that another Chinese competitor, X Square Robot, had also been excluded from the rankings. Although X Square Robot had previously ranked fourth, its exclusion appeared less contentious than the removal of Spirit AI, which had captured public imagination with its dramatic rise to the top. The incident raised questions about the reliability of such benchmarks, especially given that RoboArena is co-developed by Nvidia and academic institutions such as Stanford University and the University of California, Berkeley. The situation underscores broader concerns about the integrity of AI performance assessments, particularly in competitive environments where stakes are high. With both U.S. and Chinese firms investing heavily in physical AI research, the pressure to achieve superior results has led to increased scrutiny of how these achievements are measured and validated. The RoboArena case exemplifies the challenges of ensuring fairness and transparency in an evolving technological landscape. As the debate continues, industry experts and regulators are likely to examine the implications of this episode for future benchmarking practices. For now, the focus remains on the technical and ethical dimensions of the issue, with no immediate indication of legal action or formal sanctions against the involved parties. The outcome of this situation will undoubtedly shape how future AI advancements are evaluated and how the global community approaches the challenge of measuring progress in embodied intelligence.

1 reports

South China Morning Post logoSouth China Morning PostIndependentCenter13 hr. ago
Has a Chinese physical AI start-up manipulated a global ranking to beat Nvidia?

A Chinese AI startup, Spirit AI, briefly claimed the top spot in a global robotics benchmark called RoboArena, surpassing Nvidia. The achievement sparked controversy due to concerns about benchmark integrity. Shortly after, the benchmark's creators revised their evaluation methods and removed Spirit AI and several other models from the rankings, citing 'benchmark hacking.' The incident highlights ongoing tensions between U.S. and Chinese tech firms in AI development and raises questions about the reliability of such benchmarks. The RoboArena, developed with Nvidia and academic institutions, assesses how well robot software translates digital commands into physical actions.

Bias read (Center): The article presents a balanced account of the controversy surrounding the benchmark results without overtly favoring either the U.S. or Chinese entities. It focuses on the technical and procedural aspects of the benchmarking process rather than taking a clear ideological stance. While the broader U

How each side covered it

The same event, grouped by the political lean of the outlets covering it.

How each side covered it

Support independent, bias-aware news and unlock the social pulse, community voting, and every other Supporter feature.

Become a Supporter

Covered around the world

The same event as reported in other countries.

Covered around the world

Support independent, bias-aware news and unlock the social pulse, community voting, and every other Supporter feature.

Become a Supporter

Claims check

Key factual claims, and how many sources assert vs dispute each.

Claims check

Support independent, bias-aware news and unlock the social pulse, community voting, and every other Supporter feature.

Become a Supporter

Keep the news honest.

ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €4/month.

Become a Supporter

Related stories