Executive Summary
A new study on synthetic data for magnetic flux leakage inspection tackles a real constraint: rare defect types starve machine-learning classifiers of training samples. For operators buying in-line inspection, the decision is not whether synthetic data works in a lab but how to specify, validate and audit it inside the integrity chain. We set out the engineering trade-offs, the common procurement mistakes, and the acceptance criteria that keep a model-trained classifier honest against API 1163 and real dig verification.
A model-training problem lands in the integrity chain
The September 2026 issue of the Journal of Pipeline Science and Engineering (Volume 6, Issue 5) carries a paper by Jie Yang, Jialiang Xie, Xiaochen Gao, Kuan Fu, Jianjun Zhu and Jianli Wang on using synthetic data to improve defect classification in magnetic flux leakage (MFL) inspection when real defect samples are scarce. The core problem is one every integrity engineer recognises. MFL tools return abundant signals for general corrosion, and very few clean, labelled examples of the feature types that actually drive decisions to cut a pipe: mechanical damage, dent-with-metal-loss, narrow axial grooving, or anomalies at girth and seam welds.
Machine-learning classifiers are only as good as the distribution they were trained on. When a feature class is rare in the training set, the model tends to under-call it, over-call the majority class, or collapse the two together. Synthetic data – generated from physics models or from data-driven generative networks – is an attempt to fill those empty bins without waiting years for enough real examples to accumulate. That is the technical idea. The commercial and integrity question, and the one this article addresses, is how an operator buying in-line inspection (ILI) should treat a classifier that has been trained partly on data no dig ever produced.
Why the classifier call sits on your risk register
MFL classification is not an academic label. It feeds directly into fitness-for-service assessment and the excavation programme. A feature called “general internal corrosion” is assessed against DNV-RP-F101 or ASME B31G with a metal-loss interaction rule. The same signal called “dent with associated metal loss” or “gouge” is a mechanical-damage assessment with a different, usually more conservative, treatment and a much shorter time-to-action. Misclassification does not just move a number; it moves the feature into the wrong assessment pathway and the wrong response window.
The consequence structure is asymmetric. The feature types poorly represented in real training data are frequently the ones with the highest failure consequence. Cracking and crack-field detection, mechanical damage, and complex interacting features are precisely where an under-populated training set produces the weakest classifier, and precisely where a missed or downgraded call is most expensive. For subsea lines, where the excavation cost of a single verification dig runs to a different order of magnitude than an onshore bell-hole, the classifier’s confidence on rare features has a direct bearing on the dig budget under a DNV-RP-F116 integrity-management plan.
So the decision in front of a survey or integrity manager is not “is synthetic data clever”. It is whether adopting a model trained on partly synthetic data improves the probability of correctly identifying the features that matter, without quietly degrading trust in the reported call. Those are separable questions, and a good procurement process keeps them separate.
What actually drives the choice: two families of synthetic data
Synthetic MFL data comes from two very different engines, and the trade-off between them is the heart of the decision.
Physics-based generation solves the magnetostatic field problem directly. You define a defect geometry in a pipe wall, model the magnetiser bringing the wall toward saturation, solve the nonlinear B-H field with a finite-element or finite-difference method, and compute the leakage field at the sensor standoff. The strength of this route is controllability and coverage. You can generate any geometry you like, including feature types you have never physically encountered, and you know the ground-truth dimensions exactly because you specified them. The weakness is the sim-to-real gap. A physics model is only as faithful as its inputs: the material B-H curve, the achieved saturation level, sensor lift-off, remanent magnetism from a previous run, and – the one most often skipped – velocity effects. Real MFL signals from a gas line carry the imprint of speed excursions and dynamic magnetisation that a static solver never sees.
Data-driven generation trains a generative model – a GAN, variational autoencoder, or diffusion model – on real MFL signals and samples new examples from it. The strength here is realism: the synthetic signals inherit the noise statistics, sensor-channel behaviour and clutter of genuine field data. The weakness is coverage and control. A generative model cannot invent a class it never saw in its seed data, so it does little for the rarest features unless those already exist in some quantity. It can suffer mode collapse, and it can produce physically implausible signals that a classifier will nonetheless learn from, importing artefacts that have no counterpart in a real pipe.
The practical conclusion is that these two engines fail in opposite directions. Physics-based data covers rare geometries but carries a domain gap. Generative data closes the realism gap but under-covers the rare classes that started the problem. Hybrid approaches – physics-informed generative models, or domain randomisation over the uncertain physical parameters – exist to split the difference, and they are where the more serious work is heading. For a buyer, the point is to ask which engine produced the training data and to understand which failure mode you are therefore exposed to.
None of this changes the standing requirement that the reported calls be verified against real excavations. Synthetic data is a training-set augmentation, in the same way that a high-resolution optical survey is a supporting input rather than a substitute for the assessment itself. The same discipline we argued for when treating optical 3D scanning as a visualisation aid rather than an NDT method applies here: a tool that improves your picture of a feature does not, on its own, discharge the verification duty.
Where buyers get this wrong
1. Accepting accuracy measured on synthetic data
The single most damaging mistake is a validation set drawn from the same synthetic distribution as the training set. A classifier will always look excellent when tested on data generated by the same process that taught it. That number is meaningless as a field performance predictor. The only validation that counts is against a held-out set of real, dig-confirmed features the model never saw in training. If the vendor’s headline accuracy figure cannot be traced to a real held-out test set, treat it as a lab result, not a performance guarantee.
2. Letting synthetic data soften the dig programme
Synthetic augmentation improves the classifier; it does not add real knowledge about your pipe. The verification excavation programme required under API 1163 exists to close the loop between reported and actual condition, and it is exactly as necessary after synthetic augmentation as before. If anything, a classifier trained partly on synthetic data warrants a more deliberate dig plan on the rare classes, because those are the calls with the least real-world grounding. Sizing the dig programme down on the strength of a synthetic accuracy claim inverts the logic entirely.
3. Ignoring the magnetiser and the tool speed
A physics-based training set generated at a single saturation level, one sensor standoff, and zero velocity will teach a classifier to recognise idealised features. Real runs vary in wall thickness, achieved magnetisation, lift-off over internal deposits, and – in gas lines especially – tool speed. If the synthetic data does not span the operating envelope of the actual run, the classifier meets conditions in the field it was never shown. Ask what parameter ranges the synthetic generation covered and compare them against the run conditions of your line.
4. Buying the point estimate without the confidence and the POD
A classification result without a confidence statement is half a result. What integrity assessment needs is probability of identification (POI) and probability of detection (POD) per feature class, not a single blended accuracy across all features. A model can post 95% overall accuracy while missing half the mechanical-damage features, because that class is a small fraction of the total. Demand a per-class confusion matrix. The diagonal for your high-consequence classes is the number that governs risk, and a global average will hide a weak one.
5. Accepting an opaque training distribution
If the vendor will not state the ratio of synthetic to real data, or the per-class sample counts, you cannot judge where the model is well-supported and where it is extrapolating. This is not a demand for proprietary model internals. It is a request for the composition of the training and validation sets – the same transparency any qualified ILI process should already provide under a proper API 1163 evidence trail. An opaque distribution means an unauditable classifier, and an unauditable classifier has no place in an integrity decision.
How to decide, and what to write into the specification
The route through this is to treat synthetic-data augmentation as a qualification question and hold it to the standards that already govern ILI performance. Concrete steps:
- Anchor acceptance in API 1163. Require the classifier’s performance to be demonstrated at the applicable validation level against real, dig-confirmed features held out of training. State plainly in the ITT that synthetic-only validation is not acceptable evidence.
- Require a per-class confusion matrix on real data. Ask for POI and POD reported separately for each feature class in the Pipeline Operators Forum (POF) classification scheme, not a single blended figure. Set a minimum acceptable detection rate for the integrity-critical classes – mechanical damage, dent-with-metal-loss, axial grooving, weld anomalies – and make it a pass/fail gate.
- Disclose the training composition. Require the synthetic-to-real ratio and the per-class real sample count. Where a class is supported almost entirely by synthetic data, flag it for enhanced verification rather than accepting the call at face value.
- Interrogate the generation engine. For physics-based data, ask for the B-H curve source, achieved saturation level, sensor standoff and lift-off assumptions, and the velocity range modelled, and check those against your run conditions. For generative data, ask what physical-plausibility screening was applied and how coverage of rare classes was demonstrated rather than assumed.
- Size the dig programme to the weak classes. Prioritise excavations on the rare, high-consequence calls where synthetic augmentation did the most work. Use the results to update the model’s real-data foundation for the next run, so each campaign narrows the domain gap.
- Track the false-negative rate on critical features across runs. A missed mechanical-damage call is the failure mode that matters. Monitor it run-on-run as a standing KPI, and treat any upward drift as a trigger to re-examine the training distribution.
- Tie it to the wider integrity framework. For subsea assets, fold the classifier’s per-class performance into the DNV-RP-F116 integrity-management case and carry the sizing outputs into DNV-RP-F101 corrosion assessment with honest uncertainty, not a spuriously precise number inherited from synthetic ground truth.
Handled this way, synthetic data earns its place. It is a legitimate answer to a genuine scarcity problem, and used properly it can lift identification rates on exactly the features that starve conventional training. The failure mode is not the technique. It is buying a confident number that was never tested against a real pipe, and letting it displace the verification discipline that the integrity chain depends on. Keep the classifier honest against dig data, demand the per-class evidence, and the synthetic-training question resolves into an ordinary qualification exercise rather than an act of faith.
Based on: Synthetic data for MFL pipeline inspection defect classification
Published by
Geospatial Intelligence Group
GIS, Remote Sensing & Digital Twins
A specialist group focused on geospatial data management, GIS infrastructure, remote sensing applications, and digital twin implementation for asset lifecycle management.
Offshore Geomatics Foundation Certificate · October 2026
Turn reading into certified competence
Decisions like this one are easier to defend when the fundamentals behind them are certified, not assumed.
Early-access list gets 40% off at launch. No spam – the waitlist is only ever used for the certificate. What the certificate covers →