This piece was originally published on LinkedIn by Callum Reid, founder and CEO of Volustor.
It’s a question that keeps coming up in AI circles, and for good reason.
The current model of scaling AI has relied on ever larger quantities of text, images and video, much of it gathered from the public internet. That pool may be vast, but it is not infinite.
As Ilya Sutskever put it, there is only one internet. Epoch AI has estimated that, if current trends continue, language models could fully utilise the effective stock of public, human-generated text sometime between 2026 and 2032.
Have we hit peak data?
So, have we reached “peak data”?
Only if we have a narrow idea of what data can be.
We often imagine AI data as a mineral deposit: something already out there, waiting to be mined. Newspaper articles, Reddit forums, stock libraries and online archives. Eventually, the thinking goes, we will dig up everything useful.
But that assumes all meaningful data has already been created — that whatever mattered was already written, photographed or documented.
Reality itself is a far richer source of data than anything we have mined so far.

Consider the photograph. It flattens the world into a plane, seen from one point of view and within the narrow band of colour our eyes can register. It is a remarkable technology, but still fundamentally a nineteenth-century idea: a machine built to satisfy the human eye, not to describe everything that is there.
The future of AI requires us to think beyond the mine.
Beyond the mine: measuring reality directly
Google DeepMind’s AlphaEarth Foundations offers a clear signal. It combines multiple forms of Earth-observation data into a compact, learned representation of the planet. Each 10-by-10-metre area is represented by a 64-component embedding derived from multiple sensors and observations over time.
The result is not simply another image. It is a machine-readable record of measured reality — something that can be searched, compared, monitored and reassessed.
If AI is to become more useful, its future may depend on us creating richer data, not merely scraping more of what already exists.
What 3D scanning actually captures
That is where 3D scanning enters the conversation.
For anyone outside the field, 3D scanning captures a real object, space or person as measured spatial data, using techniques such as photogrammetry, LiDAR and time of flight.
I have spent much of my career in and around 3D scanning, and one thing you learn quickly is that a scan is closer to an observation than a finished representation. It attempts to record what was actually there, beyond a single image or point of view.
A capture might combine thousands of photographs with LiDAR measurements, depth maps, calibration references, colour charts, lens metadata, camera positions, polarisation passes, drone imagery, laser scans and material references.
That data might later become a mesh, a point cloud, a Gaussian splat, a NeRF or some form of representation that does not yet exist.
But the output is not the important part. The act of observation is.
Recording more than the eye can see
People who scan for a living work constantly against the limits of their equipment. We measure shape and geometry, but also colour, reflectance and material behaviour. We observe from the ground and from the air. Increasingly, we also capture at wavelengths the human eye cannot see.
The aim is not simply to reproduce how the world appears to us, but to record more of what is actually there.
This is why 3D scanning connects so directly to where AI is heading: towards data that is not constrained by the human point of view.
The next datasets are waiting to be created
We may be approaching the limits of public internet data. But we are nowhere near the limits of what can be observed, measured and understood.
The next great datasets do not exist online, waiting to be scraped. They are waiting to be created.
And in many industries, that process has already begun.
Movie productions, manufacturers, retailers, architects and robotics teams are generating increasingly rich spatial records of the physical world — records that extend far beyond conventional text, images and video.
Those of us working in 3D scanning, volumetric capture and spatial computing already understand something important:
What we capture today will become far more valuable tomorrow.
Frequently asked questions
Only in the narrow sense. Epoch AI estimates that language models could fully utilise the effective stock of public, human-generated internet text sometime between 2026 and 2032. But that limit applies to data that already exists online. The physical world — objects, spaces, materials, people — remains largely unmeasured, and it is a far richer source of data than anything scraped so far.
3D scanning captures a real object, space or person as measured spatial data, using techniques such as photogrammetry, LiDAR and time of flight. A single capture can combine thousands of photographs with laser measurements, depth maps, colour references and drone imagery, which can later be processed into a mesh, a point cloud, a Gaussian splat, a NeRF — or formats that do not exist yet.
Because it records reality beyond the human point of view. A scan measures geometry, colour, reflectance and material behaviour — including wavelengths the eye cannot see — rather than flattening the world into a single image. As AI moves towards understanding physical reality, spatial records like these become the datasets that cannot be scraped, only created.
If you’re sitting on spatial capture data that should be doing more than it currently is… let’s talk.
Related reading: In 2026, 3D scans will start being treated as a medium in their own right.