ON
← Volver al feed
¿Ha manipulado una start-up china de IA física un ranking global para vencer a Nvidia?
HK🏛️ PolíticaCentrohace 13 h

¿Ha manipulado una start-up china de IA física un ranking global para vencer a Nvidia?

Una startup china de IA, Spirit AI, se adjudicó brevemente el primer lugar en un benchmark global de robótica llamado RoboArena, superando a Nvidia. El logro provocó controversia debido a las preocupaciones sobre la integridad del benchmark. Poco después, los creadores del benchmark revisaron sus métodos de evaluación y eliminaron a Spirit AI y varios otros modelos de la clasificación, citando el 'hackeo de benchmark'. El incidente destaca las tensiones en curso entre las empresas tecnológicas estadounidenses y chinas en el desarrollo de IA y plantea preguntas sobre la confiabilidad de dichos benchmarks. El RoboArena, desarrollado con Nvidia e instituciones académicas, evalúa qué tan bien el software de robots traduce comandos digitales en acciones físicas.

A Chinese physical AI start-up has sparked controversy after briefly surpassing Nvidia in a global benchmark for robotic intelligence, only for its achievement to be rescinded following accusations of manipulating the evaluation process. The incident highlights the growing tensions between U.S. and Chinese firms in the race to advance artificial intelligence technologies, particularly in the field of embodied intelligence, where robots must perform tasks in unpredictable environments. The dispute began in June when Spirit AI, a startup based in Hangzhou, Zhejiang province, unveiled its latest model, Spirit v1.6, at the start of the month. The company claimed the model achieved the highest score on RoboArena, a widely recognized benchmark for assessing how well AI-driven robots can execute complex physical tasks. This result temporarily placed Spirit AI ahead of Nvidia, a leading American technology corporation known for its expertise in AI hardware and software. The benchmark, which Spirit AI described as “the Olympics of embodied intelligence in North America,” drew considerable attention due to its high-profile participants and the prestige associated with winning. However, the success was short-lived. Within days, the RoboArena team announced a major revision to its evaluation criteria and removed multiple models, including Spirit v1.6, from the official rankings. According to statements made by Pranav Atreya, a key contributor to the project and a doctoral candidate at the University of California, Berkeley, the decision followed an investigation into potential violations of the benchmark's rules. Atreya noted that the team had “retroactively removed evaluations from organizations who [it] found to be engaging in benchmark manipulation.” While he did not specify which entities were affected, the removal of Spirit AI’s model suggested that the startup had been implicated in some form of unethical behavior. The controversy surrounding RoboArena intensified further when it was revealed that another Chinese competitor, X Square Robot, had also been excluded from the rankings. Although X Square Robot had previously ranked fourth, its exclusion appeared less contentious than the removal of Spirit AI, which had captured public imagination with its dramatic rise to the top. The incident raised questions about the reliability of such benchmarks, especially given that RoboArena is co-developed by Nvidia and academic institutions such as Stanford University and the University of California, Berkeley. The situation underscores broader concerns about the integrity of AI performance assessments, particularly in competitive environments where stakes are high. With both U.S. and Chinese firms investing heavily in physical AI research, the pressure to achieve superior results has led to increased scrutiny of how these achievements are measured and validated. The RoboArena case exemplifies the challenges of ensuring fairness and transparency in an evolving technological landscape. As the debate continues, industry experts and regulators are likely to examine the implications of this episode for future benchmarking practices. For now, the focus remains on the technical and ethical dimensions of the issue, with no immediate indication of legal action or formal sanctions against the involved parties. The outcome of this situation will undoubtedly shape how future AI advancements are evaluated and how the global community approaches the challenge of measuring progress in embodied intelligence.

1 informaciones

South China Morning Post logoSouth China Morning PostIndependienteCentrohace 13 h
¿Ha manipulado una start-up china de IA física un ranking global para vencer a Nvidia?

Una startup china de IA, Spirit AI, se adjudicó brevemente el primer lugar en un benchmark global de robótica llamado RoboArena, superando a Nvidia. El logro provocó controversia debido a las preocupaciones sobre la integridad del benchmark. Poco después, los creadores del benchmark revisaron sus métodos de evaluación y eliminaron a Spirit AI y varios otros modelos de la clasificación, citando el 'hackeo de benchmark'. El incidente destaca las tensiones en curso entre las empresas tecnológicas estadounidenses y chinas en el desarrollo de IA y plantea preguntas sobre la confiabilidad de dichos benchmarks. El RoboArena, desarrollado con Nvidia e instituciones académicas, evalúa qué tan bien el software de robots traduce comandos digitales en acciones físicas.

Lectura del sesgo (Centro): El artículo presenta una explicación equilibrada de la controversia que rodea los resultados de los índices de referencia sin favorecer abiertamente a las entidades estadounidenses o chinas. Se centra en los aspectos técnicos y de procedimiento del proceso de evaluación comparativa en lugar de adoptar una postura ideológica clara.

Cómo lo cubrió cada lado

El mismo suceso, agrupado por la inclinación política de los medios que lo cubren.

Cómo lo cubrió cada lado

Apoya noticias independientes y conscientes del sesgo y desbloquea el pulso social, el voto de la comunidad y todas las demás funciones para suscriptores.

Hazte suscriptor

Cobertura en el mundo

El mismo suceso según se informó en otros países.

Cobertura en el mundo

Apoya noticias independientes y conscientes del sesgo y desbloquea el pulso social, el voto de la comunidad y todas las demás funciones para suscriptores.

Hazte suscriptor

Verificación de afirmaciones

Las principales afirmaciones fácticas y cuántas fuentes las respaldan o las rebaten.

Verificación de afirmaciones

Apoya noticias independientes y conscientes del sesgo y desbloquea el pulso social, el voto de la comunidad y todas las demás funciones para suscriptores.

Hazte suscriptor

Mantengamos las noticias honestas.

ObjectiveNews se financia con los lectores y no tiene anuncios: te mostramos el sesgo en lugar de ocultarlo. Apoya el periodismo independiente por 4 €/mes.

Hazte suscriptor

Historias relacionadas