Deep learning models usually need hundreds or thousands of labeled examples to learn how to identify specific archaeological features. While this is possible for common and well studied discoveries, rare or previously unknown features are much harder to collect enough data for.
A research team has found a solution by creating realistic artificial examples using code and training the model on those instead of relying only on real data.
The core problem
Most archaeological deep learning projects eventually face the same problem. AI models need enough labeled examples to understand the different ways an object can appear, but many archaeological features simply don’t occur often enough to build one.
Common methods like flipping, stretching, or blurring existing images can create more training data, but they have limits. When there are fewer than about 100 real examples, these techniques can cause the model to memorize the available images instead of actually learning the object’s shape and features.
Because of this, deep learning has stayed out of reach for a meaningful share of archaeological work, not because the tools don’t work, but because the raw material to train them on was never there to begin with.
Procedural generation
The team used procedural generation: code that automatically creates realistic synthetic archaeological objects and places them within real terrain data. The process also generates corresponding training labels, eliminating the need for manual annotation. Similar techniques have been used in other fields, including healthcare, where synthetic data can help researchers train models without sharing sensitive patient records.
To test it, the team turned to a real case: 12 unusual earthen structures in a Louisiana forest, rare enough that no traditional training dataset could be built around them. They compared three approaches to generating training data instead.
The first built complete synthetic structures from scratch, placing them in realistic locations while avoiding steep slopes and drainage areas, and adding rough circular edges, an inner depression, and a secondary feature. These artificial structures were embedded directly into real terrain data, with training labels generated automatically alongside them.

The second approach stripped things down to basic circular shapes with none of the extra detail, testing whether simpler geometry alone could still teach the model to recognize the target features.
The third used hundreds of real examples of a related feature from another region, mathematically reshaping them to resemble the Louisiana structures instead of simulating anything from scratch.
Each dataset trained its own model under identical settings, then all three were tested against the same real-world data to see which approach performed best.
Results
The results revealed a tradeoff between precision and detection. Models trained on detailed synthetic objects were the most accurate overall but missed some genuine examples. Models trained on simpler synthetic objects detected all of the real structures, but also generated more false positives that required manual review. A model trained on modified real data produced intermediate results.
Overall, the researchers found that the more closely the synthetic objects matched the shape of the real archaeological features, the more precise the model became. Simpler, more general synthetic objects helped the model find more potential examples, but they also produced more false detections that researchers had to remove manually.

Improving results
Because all three methods produced false positives, mainly natural landforms that looked similar to the target structures, the researchers added an automated filtering step to improve the results.
The system analyzed the elevation and curvature of each detected feature to distinguish natural landforms from human-made structures. Natural dome-shaped landforms had a different curvature pattern than the archaeological structures.

This filtering step removed many false positives, reducing the need for manual review and improving the model’s performance.
Field validation
To confirm that the models were detecting real archaeological features rather than random patterns, the researchers visited some of the highest-confidence locations in the field and excavated selected sites.
The structures were confirmed to be real, although their findings suggested they had a different historical purpose than the team originally expected. This fieldwork provided important evidence that the synthetic data approach could successfully identify genuine archaeological features rather than false patterns created by the simulations.
The tradeoff
There are still some limitations to the approach. The more generic synthetic-object method produced a large number of false positives, with thousands of potential candidates requiring automated and manual filtering before reaching a workable shortlist. The researchers argue that this cleanup process is still much faster than conducting a full manual survey of the same area, but the method is not completely hands-off.
The training datasets in this comparison were also deliberately kept relatively small, with fewer than 400 synthetic objects per method to ensure a fair comparison. However, procedural generation could produce training sets many times larger in a similarly short amount of time, which could potentially improve precision even further. The study did not test whether larger datasets would improve precision.
Overall, the results suggest that synthetic data can make deep learning practical for rare archaeological features, while still requiring some filtering and human review.
Peck, K., Gravel-Miguel, C., Snitker, G., & Helmer, M. “Using Simulated Training Data to Locate Archaeological Sites with Machine Learning.” Advances in Archaeological Practice, Vol. 14, Issue 2 (2026), pp. 179–197. doi.org/10.1017/aap.2025.10130
Open access under CC BY 4.0. Code available via GitHub and Zenodo (doi.org/10.5281/zenodo.17082763).