AI models operate on a particular domain and data type. Everybody has interacted with a chat model. Models like ChatGPT are pretrained on large, unstructured bodies of text (starting with the entire internet), without supervision, using something called a pretext task. For ChatGPT it’s next-token prediction: trying to predict the next fragment of a word, the next word in a sentence. That turns out to be a really useful task for getting a large model to learn the underlying structure of data. For that reason, domain matters. If you want to generate Shakespearean prose, you train on the complete works of Shakespeare. If you want code, you train on large bodies of code.
To date, there haven’t been any large-scale models developed for industrial X-ray CT data. That’s what we’ve built at Lumafield: the first large-scale model for understanding this uniquely rich type of volumetric data.
This has been attempted with medical data, and it has obviously been done with natural images, but neither maps effectively to industrial CT. It’s similar to the transfer gap between Shakespearean sonnets and Python programs. They’re both words, but they come from very different domains and the rules are very different. Industrial CT images form in a very different way. They’re not three-channel color images, they have different bit depths, and the image formation physics is also different. A natural image forms by visible light reflecting off materials, scattering, absorbing, refracting through a lens onto an image sensor.
Industrial CT, meanwhile, is a computational imaging process. In a reflection image, a pixel’s value communicates something about the surface it hit. In an X-ray image, on the other hand, the value of a pixel describes something about depth in space. That limits what models trained on natural images can extract from CT data, and it’s why we set out to train a model purpose-built for extracting high-quality features on this domain.
This is the dataset behind our foundation model. It’s unique for three reasons:
That domain gap is why we trained from scratch. To understand what the model actually produces, you need one concept: the embedding. One useful, somewhat simplified metric of a foundation model is whether it encodes similar samples to similar locations in embedding space. If you show a population of dogs and cats to a classic vision model, the dogs should form one cluster, the cats another. Maybe a coyote lands closer to the dogs and a tiger closer to the cats. There’s meaning to proximity, distance, and direction in the space.
A well-formed embedding has directions that matter for a task. In early face models there’d be a direction for hair color, directions for emotional state, happy versus sad, wearing glasses versus not. A good embedding’s constituent axes map to meaningful variation in the dataset. Ours currently runs 798 dimensions.
Part of embedding quality is what the model learns to ignore. Show it the same image at different noise levels. Does it embed to the same location, or does noise spread it in a clear direction? Our model does see data quality, blurrier versus sharper shows up, which is itself useful.
Two subtly different claims are easy to conflate here. One is that individual dimensions carry specific meaning, an axis for battery chemistry, a direction for missing screws. That isn’t how our model currently works, and it doesn’t need to. The other is that scans with different structure land in different places in the space. That one is definitely true, and that’s the premise of the model.
The best way to judge a model like this is to look through its eyes at real populations of parts. We’ve demonstrated that at scale three times, with phones, connectors, and batteries.
Our dataset of several hundred aftermarket iPhone 7s clusters into regions of self-similarity. These phones were sold with two main logic board configurations, one using a Qualcomm chip and one using an Intel chip. The different packages and different pinouts manifest as different solder patterns in one region of the main logic board under the A10 processor. The model distinguishes the two chipsets on its own.
Because these are aftermarket phones, many have had their batteries replaced, and clusters of different battery construction and different battery vendors form as well: one region where the chemistry has a very prominent tab, another where the polarity of the battery construction is different. One phone is visibly missing its screws. Someone must’ve opened it up, replaced the battery, and didn’t put the screws back in.
What’s powerful is that the model interprets this data at scale and finds hidden patterns that correlate with useful, insightful properties of manufactured goods. These are sources of variation a product designer or refurb manager may not have considered would even exist.
The phones show the model discovering variety nobody cataloged. Our connectors show it drawing a line that matters commercially. Look at real vs. fake through the model’s eyes, and you get a really nice breakdown of the connector data: all of one type clusters on one side, all of the other type on the other. Look at what varies, and one population has less dense gaskets than the other, and a component sits in a slightly different position.
We scanned 1,496 connectors across eight part numbers on a Triton: genuine units from authorized distributors, and open-market equivalents of the same part numbers from AliExpress. The embedding view holds up, and when we go back into the volumes to ask why the populations separate, the answers are physical. We were able to measure differences in material choice. Denser but more transparent to X-rays is a revealing combination: glass fill raises attenuation strongly, so removing the glass and swapping to a heavier base polymer produces exactly this signature. Our hypothesis is that an unfilled resin like PBT was substituted to save on tool wear and molding difficulty.
The consequence comes later, in the field. Without glass fill, the housing has less strength and dimensional stability when subjected to temperature cycling and vibration. Porosity tells the same story. The median open-market part has about 5x more porosity than the median genuine part. On a single part number, the model classified genuine vs. open market with 98.8% accuracy.
Last year we scanned 1,054 batteries: 100 samples from each of 10 manufacturers, all the same cylindrical 18650 form factor. We included OEMs, mid-tier reputable brands, and pure grey-market brands.
When we passed the data through the foundation model and looked at the scans through its eyes, we found clear separation. The model could distinguish cells from different manufacturers. But there was one case where the model perfectly failed to distinguish cells from two different vendors. When we looked closer at why, it turned out one vendor was an OEM and the second was a rewrap vendor, a company that purchases cells from an OEM, rewraps them, white-labels them, and sells them under a different brand. The fact that the model couldn’t structurally distinguish the cells cleanly indicates they’re actually the same population of cells. The failure was the discovery. That's a provenance risk you could never see before, visible in the parts themselves.
There’s a bigger argument here than any single case. In a lot of electronics inspection there’s an AXI system taking 2D X-ray images, or maybe even a CT, and the bottleneck is putting eyes on those images.
Human eyes that can process those images and extract insights and decisions are exceedingly scarce, from both a labor availability and training perspective. The only thing that scales with the volume of data produced by vision systems and in-line systems like Mars is machine perception, a foundation model tuned to understand exactly those images.
Quality Agent and AIM build on the foundation model. Together, the three form a framework for automating and extending quality in manufacturing, and mark a step towards our goal of automating manufacturing. They address two key bottlenecks we’ve observed time and time again collaborating with hundreds of customers across every major industry, and we’re confident from years working in this space that these are the real bottlenecks that hold back scaling.
Lumafield Quality Agent derives its perception from this model. It joins a deep understanding of part quality with process control knowledge to autonomously close the loop.
And it all starts from the same place: a model that learns from a kind of data no one else has.