Beijing’s first humanoid robot data training center, once considered a benchmark for embodied intelligence infrastructure, recently underwent adjustments.
On August 28, media reported that the humanoid robot data training center, located in Shougang Park, Shijingshan District, Beijing, had ceased operations, with the core equipment and technology provider, Realman, withdrawing.
Subsequently, Shijingshan District offered a different explanation: after the original partner of Phase I relocated, the center is undergoing technical capability upgrades and will focus on 4D Gaussian light fields, world models, and other technical routes in the future, forming synergy with Phases II and III.
Therefore, rather than interpreting this as a data center suddenly “collapsing,” a more noteworthy issue is that after more than a year of operation, Beijing’s earliest humanoid robot data collection centers are proactively changing course.
This precisely exposes an increasingly real problem in the embodied intelligence data industry: building a data center is not difficult; the difficulty lies in who will pay for the data in the long term.
There are many data centers in China, but where is the real demand?
Phase I of the Shijingshan Humanoid Robot Data Training Center was launched in March 2025, covering an area of approximately 3,000 square meters and deploying 108 robots. It established ten scenarios, including home healthcare, special operations, new retail, automobile assembly, and 3C electronics factories, initially planning to produce over one million high-quality multimodal data points annually.

Its expansion has been rapid since then.
In October 2025, Phase II, covering approximately 6,000 square meters, began operation; by May of this year, Phase III of the Shijingshan Data Training Center had gathered nearly 270 robots, with an officially disclosed annual data production capacity approaching 20 million data points.
Similar projects have emerged in large numbers across China.
Over the past two years, data collection centers have become an important tool for many regions to develop embodied intelligent industries. The logic is not difficult to understand: individual startups invest heavily in purchasing robots, renting space, building remote operating systems, and recruiting data collection personnel. Public training grounds can spread these costs while simultaneously undertaking functions such as testing and verification, talent training, standard setting, and industry investment promotion.
The problem arises in the next step.
Data centers can be built through project investment, but their operation ultimately requires real data support.
Embodied intelligence data currently possesses strong “private domain” attributes.
Different companies use different robot bodies, sensors, dexterous hands, control frequencies, and model architectures; even when performing the same “folding clothes” action, data collected by one company is difficult to be used for another company completely cost-free.
This means that while robotics companies certainly need data, they may not be willing to continuously purchase general-purpose data.
Thus, a potential contradiction gradually emerges: local governments building training centers focus more on industrial public service capabilities, enterprise agglomeration, and infrastructure supply; robotics companies, on the other hand, focus on whether data can directly improve model capabilities and whether the cost of acquiring this data is lower than self-collection.
These two sets of goals overlap but are unlikely to naturally align.
This is where the risk of so-called “data real estate” arises: sites, robots, remote-controlled devices, and computing power can all be quantified, but the true value of data is difficult to measure by area, number of devices, and collection time.
Shift from “How much data there is” to “Is the data useful?”
Realman’s explanation this time is very straightforward.
According to the company, in the past, remote data collection in training environments was conducted with controllable environmental parameters, stable lighting, and relatively fixed scenarios, resulting in overly “clean” data. However, when robots actually enter factories, warehouses, or homes, they face material deviations, lighting changes, ground disturbances, and numerous abnormal operating conditions, which training environments struggle to fully reproduce.
Therefore, Realman is shifting from centralized remote operation in the laboratory to remote operation of robots in real-world scenarios, allowing data to continuously flow back with actual operations.
This represents a change in the logic of embodied data.
Early in the industry, the focus was on “whether there was data.” At that time, real-world data was scarce, and large-scale remote operation training environments could quickly supplement basic motion data, which was still valuable for model cold starts and basic skill learning.
However, as robots are increasingly being applied, the question is shifting to “which data is useful.”
Compared to repeatedly collecting large amounts of standard motion data, failures, error corrections, disturbances, and long-tail situations in real-world environments often bring greater incremental value to the model. Evaluating a data acquisition center by simply calculating collection hours, data entries, or the number of robots is increasingly failing to reflect the final training effect of the data.
Simultaneously, the data routing itself is also rapidly changing.
Beyond teleoperation of physical robots, simulation data, synthetic data, video data, internet data, and world models are all becoming new data sources for embodied models. This year, the National Data Administration has even initiated the formulation of standards such as the “High-quality dataset—Embodied intelligence—Specifications for simulated synthetic data generation and processing” and the “High-quality dataset—Embodied intelligence—Specification for data collection and model training in training bases.”
This also means that it is still difficult to predict in advance whether a particular data production method will become the final answer.
Data acquisition centers need to move from “building projects” to “running closed loops”
Therefore, the adjustment of Beijing’s first humanoid robot data training center does not necessarily mean that public data training grounds have lost their value.
What truly needs to be re-examined is its role.
In June of this year, the National Data Administration released the Implementation Plan for the Initiative to Promote the Construction of High-Quality Industrial Datasets, proposing “demand-driven, urgent-use-first, and application-verified,” emphasizing the formation of a data flywheel of “scenario-driven data, data-driven models, models-empowered applications, and applications creating value.” The document also proposes the construction of a national dataset management and service system that is “physically distributed but logically centralized.”
This may be closer to the next stage than simply building more physical training grounds.
The value of future public data collection platforms may not primarily lie in the number of square meters of space they possess or the number of robots they purchase, but rather in their ability to connect dispersed scenarios such as real factories, supermarkets, logistics, and healthcare; to establish cross-enterprise data format, quality evaluation, and authorization mechanisms; and to provide testing, evaluation, and public data services that startups cannot afford on their own.
Correspondingly, the evaluation system also needs to change.
It should shift from “how many data points were produced” to “how much data was actually used by the model”; from equipment utilization to the task success rate and generalization ability after model training; and from one-time construction investment to whether companies are willing to continuously purchase services.
Only when demanders are willing to continuously pay can data centers truly form a commercial closed loop.
In this sense, the shift in focus of Beijing’s first humanoid robot data training center provides a valuable industry example.
Over the past year, various regions have been rapidly replicating the “space + equipment + robot + data collector” training center model. As the industry enters the next stage, the question has shifted from “whether there is data infrastructure” to “how much effective data these infrastructures actually create.”
Embodied AI is still in a period of rapid technological iteration. For local industrial policies, infrastructure construction remains necessary, but more important than expanding area and equipment scale is ensuring that investment adapts to real demand.
The truly scarce resource in data centers is never buildings or robots, but a continuously operating data flywheel.



