- The timeline of an idea born in university labs
- The bottleneck nobody wants to talk about
- Why India and why now
- Winners, losers, and those watching from the sidelines
- The SHM Studio take: data as infrastructure, not as a product
- Operational implications for the Italian market
- The work in progress: unresolved issues
- Next moves: what to monitor over the next 18 months
Human Archive is a startup founded by researchers from UC Berkeley and Stanford. Its model is simple but radical: pay Indian gig economy workers to wear caps with cameras and sensory devices. The goal is to collect real-world physical data. This data is used to train robots and embodied artificial intelligence systems, so-called Physical AI .
Therefore, the project intercepts one of the most critical bottlenecks of modern AI: the scarcity of quality physical data. In fact, while language models feed on already abundant digital text, robots need to observe movements, environments, and human interactions in the real world. Furthermore, India's choice is not random: a mature digital services ecosystem and a large workforce significantly lower collection costs.
In short, Human Archive represents a relevant case study for anyone working in the AI ecosystem. We at SHM Studio we analyze it to understand where infrastructure investments are shifting in the sector and what operational implications emerge for Italian SMEs considering the adoption of advanced AI solutions.
The timeline of an idea born in university labs
Human Archive was born from the meeting between the elite academic world and the concrete needs of the robotics industry. The founders come from UC Berkeley and Stanford University. Both universities have been at the center of research on autonomous robots and reinforcement learning for years. However, the leap from theory to practice requires something that university labs cannot produce at scale: real-world physical data.
The startup has therefore structured an operating model based on the gig economy. Workers recruited in India wear caps equipped with cameras and sensors. These devices record movements, domestic environments, interactions with everyday objects, and spatial dynamics. Subsequently, the data is processed and sold to AI and robotics labs that use it to train their models.
According to reports by TechCrunch , the project is part of a global race to acquire physical training data. In fact, demand from AI labs and robotics companies is growing rapidly. Consequently, whoever controls the data collection infrastructure gains a structural competitive advantage.
The bottleneck nobody wants to talk about
Public debate on artificial intelligence often focuses on large language models. However, there is another equally strategic frontier: the Physical AI , or systems capable of acting in the physical world. Think of industrial robots, autonomous vehicles, home assistance systems.
To train these systems, data is needed that digital datasets cannot provide. Specifically, video sequences of human movements in real environments, recordings of interactions with objects, and sensory maps of domestic and work spaces are required. Therefore, collecting this data has become one of the most expensive infrastructural challenges in the entire AI sector.
According to the analyses of McKinsey , physical automation represents one of the most significant economic growth drivers of the decade. Therefore, whoever solves the physical data problem isn't just building a company: they are positioning themselves as critical infrastructure for a multi-billion dollar industry.
Why India and why now
Human Archive's geographical choice is not random. India has a mature digital services ecosystem, with gig economy platforms already operational and a workforce accustomed to structured digital tasks. Furthermore, the cost of labor allows for economically sustainable data collection margins compared to Western markets.
Similarly, timing is crucial. In 2025, there was an acceleration of investments in robotics by major global tech players. Consequently, the demand for quality physical data exploded in a context of still very limited supply. Human Archive has positioned itself precisely in this void.
Among other things, the operating model replicates – in a physical key – what companies like Scale AI have done for textual and visual data. Therefore, the industrial precedent exists and has already been validated by the markets. The difference is that collecting physical data requires physical presence, which makes geographic distribution a critical factor of competitive advantage.
Winners, losers, and those watching from the sidelines
In this scenario, very different positions emerge among market players. The short-term winners are clearly robotics labs and Physical AI companies that gain access to previously inaccessible datasets. Furthermore, the workers of the Indian gig economy win, finding a new category of paid micro-tasks.
On the contrary, the potential losers are companies that are building robotic solutions without solving the data problem. Despite this, many of these players have not yet perceived the urgency of the problem. Therefore, they risk finding themselves at a structural disadvantage compared to competitors who have invested earlier in data infrastructure.
Finally, there's a third category: those who observe without yet acting. Many Italian SMEs in manufacturing and logistics are evaluating robotic automation solutions. For this reason, understanding where the bottlenecks in the physical AI ecosystem are is strategically relevant even for those who do not operate directly in the tech sector.
The SHM Studio take: data as infrastructure, not as a product
We at SHM Studio We're looking at the Human Archive case through a specific strategic lens. The point isn't the startup itself. The point is the paradigm shift it represents: physical data is becoming critical infrastructure, just as digital behavioral data was for programmatic marketing.
This has direct implications for Italian SMEs operating in applied artificial intelligence . In fact, the most advanced AI solutions — from computer vision to collaborative robotics — will increasingly depend on the quality of physical training data. Consequently, whoever controls or has privileged access to this data will have a competitive advantage that is difficult to overcome.
Furthermore, the Human Archive model suggests that global-scale data collection requires distributed architectures and partnerships with local ecosystems. Therefore, even companies that do not produce robots must start thinking in terms of data supply chain physical, not just digital.
Operational implications for the Italian market
For Italian SMEs in manufacturing, logistics, and retail, the implications are concrete. First and foremost, those considering the adoption of robotic solutions should include the quality and origin of training data in their evaluation criteria. This factor is often overlooked in purchasing analyses.
Secondly, companies that already collect physical operational data — warehouse videos, production line recordings, sensor logs from machinery — might possess strategic assets that are not yet valued. In fact, this data could become the subject of partnerships or licensing with players in the physical AI sector.
Finally, those responsible for Digital marketing and SEO in the B2B sector should monitor the evolution of the Physical AI sector as an emerging vertical. Among other things, the opportunities for organic positioning on these topics are still very open in the Italian market. Similarly, campaigns Linkedin and Google Ads on keywords related to robotics and physical AI still show relatively low levels of competition.
The work in progress: unresolved issues
The Human Archive model also raises questions that the market has not yet solved. The first concerns the privacy and informed consent . Collecting environmental data through workers wearing cameras opens up complex regulatory scenarios, especially in view of a possible European expansion. The European AI Act imposes strict constraints on the collection and processing of biometric and environmental data.
The second issue concerns the quality and representativeness of the collected data. Physical data coming primarily from Indian environments may not be sufficiently representative for robots intended to operate in European or North American contexts. Therefore, geographic diversification of data collection will be a central theme in the coming years.
Finally, the issue of sustainability of the gig model . If demand for physical data grows as expected, pressure on workers and compensation could create tensions. Despite this, there are currently no industry standards for fair compensation for this type of work.
Next moves: what to monitor over the next 18 months
Looking ahead to 2027-2028, three developments are worth noting. The first is the potential entry of major tech players — Google, Meta, Amazon — into the physical data collection market, through acquisitions or internal development. This will rapidly redefine competitive dynamics.
The second is the evolution of European regulations. In fact, the European Commission is already working on specific guidelines for embodied AI systems. Consequently, companies operating in this space will need to adapt their operating models in advance.
The third is the birth of specialized marketplaces for physical data. Similar to what happened with digital data, dedicated trading and licensing platforms are likely to emerge. Those who position themselves now — even just as informed observers — will have an advantage in understanding these opportunities. To delve deeper into the AI implications for your company, the team at SHM Studio is available for a dedicated consulting . Furthermore, on our Blog we regularly publish analyses on Web , content strategy and digital innovation for the Italian market.
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.