A Chinese physical AI start-up has sparked controversy after briefly surpassing Nvidia in a global benchmark for robotic intelligence, only for its achievement to be rescinded following accusations of manipulating the evaluation process. The incident highlights the growing tensions between U.S. and Chinese firms in the race to advance artificial intelligence technologies, particularly in the field of embodied intelligence, where robots must perform tasks in unpredictable environments. The dispute began in June when Spirit AI, a startup based in Hangzhou, Zhejiang province, unveiled its latest model, Spirit v1.6, at the start of the month. The company claimed the model achieved the highest score on RoboArena, a widely recognized benchmark for assessing how well AI-driven robots can execute complex physical tasks. This result temporarily placed Spirit AI ahead of Nvidia, a leading American technology corporation known for its expertise in AI hardware and software. The benchmark, which Spirit AI described as “the Olympics of embodied intelligence in North America,” drew considerable attention due to its high-profile participants and the prestige associated with winning. However, the success was short-lived. Within days, the RoboArena team announced a major revision to its evaluation criteria and removed multiple models, including Spirit v1.6, from the official rankings. According to statements made by Pranav Atreya, a key contributor to the project and a doctoral candidate at the University of California, Berkeley, the decision followed an investigation into potential violations of the benchmark's rules. Atreya noted that the team had “retroactively removed evaluations from organizations who [it] found to be engaging in benchmark manipulation.” While he did not specify which entities were affected, the removal of Spirit AI’s model suggested that the startup had been implicated in some form of unethical behavior. The controversy surrounding RoboArena intensified further when it was revealed that another Chinese competitor, X Square Robot, had also been excluded from the rankings. Although X Square Robot had previously ranked fourth, its exclusion appeared less contentious than the removal of Spirit AI, which had captured public imagination with its dramatic rise to the top. The incident raised questions about the reliability of such benchmarks, especially given that RoboArena is co-developed by Nvidia and academic institutions such as Stanford University and the University of California, Berkeley. The situation underscores broader concerns about the integrity of AI performance assessments, particularly in competitive environments where stakes are high. With both U.S. and Chinese firms investing heavily in physical AI research, the pressure to achieve superior results has led to increased scrutiny of how these achievements are measured and validated. The RoboArena case exemplifies the challenges of ensuring fairness and transparency in an evolving technological landscape. As the debate continues, industry experts and regulators are likely to examine the implications of this episode for future benchmarking practices. For now, the focus remains on the technical and ethical dimensions of the issue, with no immediate indication of legal action or formal sanctions against the involved parties. The outcome of this situation will undoubtedly shape how future AI advancements are evaluated and how the global community approaches the challenge of measuring progress in embodied intelligence.
★
Ohranimo novice poštene.
ObjectiveNews financirajo bralci in je brez oglasov – pristranskost vam pokažemo, ne skrijemo. Podprite neodvisno novinarstvo za 4 €/mesec.
Postani podpornik