White Paper

From Data Labelling to Deployed Detection: AI in Aviation Security

What it really takes to put an AI detection system into a regulated security environment — and why the dataset, not the model, decides the outcome.

Artificial Intelligence · Security7 min read
← All white papers

Automatically detecting prohibited items in cabin baggage is exactly the kind of problem AI should solve: high volume, repetitive, safety-critical, and constrained by human fatigue. It is also exactly the kind of environment where AI cannot be deployed casually — regulated, operationally sensitive, and intolerant of silent failure.

ClearPath Partnership worked with a technology partner specialising in AI-powered decision intelligence to develop a system for automatically identifying prohibited items in cabin baggage — providing AI development and a dedicated data labelling and annotation service. This paper shares what that programme taught us about deploying AI in security operations.

The dataset is the product

The public conversation about AI fixates on models. In applied detection work, the model is increasingly a commodity; the dataset is the differentiator. Detection performance is bounded by the quality, coverage and consistency of the labelled data the model learns from — and security imagery is hard data: cluttered bags, overlapping items, deliberate concealment.

We built a data labelling and annotation service as a first-class part of the programme, not a back-office task. That meant trained annotators with clear labelling standards, quality assurance on the annotations themselves, and a feedback loop where model errors drove targeted data collection. If a detection system is underperforming, the answer is usually in the data pipeline, not the model architecture.

Build for the regulator from day one

Aviation security is a regulated domain. Requirements around performance evidence, auditability and operational procedure are not obstacles to deployment — they are the specification. Programmes that treat compliance as a final-stage gate discover late that their architecture cannot produce the evidence regulators need. We engaged operational requirements, security standards and regulatory compliance as design inputs alongside accuracy targets.

Humans stay in the loop

The goal of AI in screening is not to remove the human — it is to make human attention count where it matters. Detection systems should be designed around the operator: clear presentation of what was flagged and why, workflows that support quick confirmation or rejection, and performance monitoring that captures operator disagreement as a signal for improvement. Systems designed to assist are adopted; systems that appear to replace judgement are resisted.

Measure operational outcomes, not model metrics

Precision and recall matter, but the metrics that decide whether a deployment succeeds are operational: throughput, false-alarm workload, operator confidence and security outcome. A model improvement that adds review workload can make the operation worse. Programme governance should put operational measures alongside model measures from the first pilot.

Key takeaways

  • Invest in labelling and annotation as a product capability — dataset quality bounds everything.
  • Treat regulatory and operational requirements as design inputs, not late-stage gates.
  • Design for the operator: AI that assists gets adopted.
  • Close the loop — model errors should drive targeted data work.
  • Judge the system by operational outcomes, not leaderboard metrics.

Where this applies beyond screening

The same pattern — curated data, regulatory-first design, human-in-the-loop operations — applies wherever AI meets a regulated or safety-critical process: infrastructure inspection, medical triage, industrial quality control. The screening lane is simply where aviation learned it first.

Closing thoughts

Gareth Wilson

“The uncomfortable engineering truth is that the dataset, not the model architecture, decides the outcome. Invest in labelling quality, ground truth and edge-case coverage first — the model improvements follow almost for free.”

Gareth Wilson · Founder and Co-CEO

Phil Moss

“Delivering AI into a regulated security environment is a programme, not an experiment. Clear acceptance criteria, staged evidence for the regulator, and honest reporting on model performance kept this deployable rather than perpetually promising.”

Phil Moss · Founder and Co-CEO

Talk to the people who did the work. If the challenges in this paper look like yours, we'd be glad to share more of what we've learned — and how it could apply to your organisation.

Get in touch