Are We Running Out of Data?

July 22, 2026

Iceberg diagram: a small tip above the waterline shows stacked documents and charts representing scraped internet data, while a much larger mass below the waterline shows a city, bridge, human head, terrain contours, robotic arm, wind turbines and ship representing the vastly larger unmeasured physical world

This piece was originally published on LinkedIn by Callum Reid, founder and CEO of Volustor.

It’s a question that keeps coming up in AI circles, and for good reason.

The current model of scaling AI has relied on ever larger quantities of text, images and video, much of it gathered from the public internet. That pool may be vast, but it is not infinite.

As Ilya Sutskever put it, there is only one internet. Epoch AI has estimated that, if current trends continue, language models could fully utilise the effective stock of public, human-generated text sometime between 2026 and 2032.

Have we hit peak data?

So, have we reached “peak data”?

Only if we have a narrow idea of what data can be.

We often imagine AI data as a mineral deposit: something already out there, waiting to be mined. Newspaper articles, Reddit forums, stock libraries and online archives. Eventually, the thinking goes, we will dig up everything useful.

But that assumes all meaningful data has already been created — that whatever mattered was already written, photographed or documented.

Reality itself is a far richer source of data than anything we have mined so far.

Iceberg diagram: a small tip above the waterline shows stacked documents and charts representing scraped internet data, while a much larger mass below the waterline shows a city, bridge, human head, terrain contours, robotic arm, wind turbines and ship representing the vastly larger unmeasured physical world
The internet text AI has already scraped is the tip of the iceberg. Everything below the waterline still has to be observed and measured.

Consider the photograph. It flattens the world into a plane, seen from one point of view and within the narrow band of colour our eyes can register. It is a remarkable technology, but still fundamentally a nineteenth-century idea: a machine built to satisfy the human eye, not to describe everything that is there.

The future of AI requires us to think beyond the mine.

Beyond the mine: measuring reality directly

Google DeepMind’s AlphaEarth Foundations offers a clear signal. It combines multiple forms of Earth-observation data into a compact, learned representation of the planet. Each 10-by-10-metre area is represented by a 64-component embedding derived from multiple sensors and observations over time.

The result is not simply another image. It is a machine-readable record of measured reality — something that can be searched, compared, monitored and reassessed.

If AI is to become more useful, its future may depend on us creating richer data, not merely scraping more of what already exists.

What 3D scanning actually captures

That is where 3D scanning enters the conversation.

For anyone outside the field, 3D scanning captures a real object, space or person as measured spatial data, using techniques such as photogrammetry, LiDAR and time of flight.

I have spent much of my career in and around 3D scanning, and one thing you learn quickly is that a scan is closer to an observation than a finished representation. It attempts to record what was actually there, beyond a single image or point of view.

A capture might combine thousands of photographs with LiDAR measurements, depth maps, calibration references, colour charts, lens metadata, camera positions, polarisation passes, drone imagery, laser scans and material references.

That data might later become a mesh, a point cloud, a Gaussian splat, a NeRF or some form of representation that does not yet exist.

But the output is not the important part. The act of observation is.

Recording more than the eye can see

People who scan for a living work constantly against the limits of their equipment. We measure shape and geometry, but also colour, reflectance and material behaviour. We observe from the ground and from the air. Increasingly, we also capture at wavelengths the human eye cannot see.

The aim is not simply to reproduce how the world appears to us, but to record more of what is actually there.

This is why 3D scanning connects so directly to where AI is heading: towards data that is not constrained by the human point of view.

The next datasets are waiting to be created

We may be approaching the limits of public internet data. But we are nowhere near the limits of what can be observed, measured and understood.

The next great datasets do not exist online, waiting to be scraped. They are waiting to be created.

And in many industries, that process has already begun.

Movie productions, manufacturers, retailers, architects and robotics teams are generating increasingly rich spatial records of the physical world — records that extend far beyond conventional text, images and video.

Those of us working in 3D scanning, volumetric capture and spatial computing already understand something important:

What we capture today will become far more valuable tomorrow.

Frequently asked questions

Is AI really running out of training data?

Only in the narrow sense. Epoch AI estimates that language models could fully utilise the effective stock of public, human-generated internet text sometime between 2026 and 2032. But that limit applies to data that already exists online. The physical world — objects, spaces, materials, people — remains largely unmeasured, and it is a far richer source of data than anything scraped so far.

What is 3D scanning?

3D scanning captures a real object, space or person as measured spatial data, using techniques such as photogrammetry, LiDAR and time of flight. A single capture can combine thousands of photographs with laser measurements, depth maps, colour references and drone imagery, which can later be processed into a mesh, a point cloud, a Gaussian splat, a NeRF — or formats that do not exist yet.

Why is 3D scan data valuable for AI?

Because it records reality beyond the human point of view. A scan measures geometry, colour, reflectance and material behaviour — including wavelengths the eye cannot see — rather than flattening the world into a single image. As AI moves towards understanding physical reality, spatial records like these become the datasets that cannot be scraped, only created.

If you’re sitting on spatial capture data that should be doing more than it currently is… let’s talk.

Related reading: In 2026, 3D scans will start being treated as a medium in their own right.

Picture of Callum Rex Reid

Callum Rex Reid

With over 15 years of experience in 3D capture technologies, and digital product innovation, last year I tansitioned from Managing Director for Europe at Visualskies Ltd to CEO of our new venture Volustor Ltd.