Designing a GPS-Denied Navigation Trial: Ground Truth, Metrics and Failure Cases

A trial that only flies easy ground on clear days proves very little. How to design a GPS-denied navigation trial that tells you something useful.

Every navigation system looks good in a video of a single flight on a clear morning over distinctive ground. That is not evidence. Whether you are evaluating a product, validating an integration or gathering material for a safety case, the value of a trial depends on its design: what question it asks, how the truth is established, what is measured, and whether the conditions that cause failures were deliberately included. This article sets out what a good GPS-denied navigation trial looks like.

Start with the decision

Before planning a single flight, write down the decision the trial is meant to inform. “Is this system good enough to act as a GPS fallback on our corridor inspections?” leads to a very different trial from “Can this system support a mission with no GPS at all?” The first cares most about transitions and short denials over familiar ground. The second cares about long denials, cold starts and every kind of terrain on the route.

From the decision, define the operating envelope the trial must cover: the altitude band, speeds, terrain types, times of day, seasons and weather that the real operation will see. Then write the pass and fail criteria while you are still neutral about the result. Criteria written after the data arrives have a way of fitting the data. If the results will support a regulatory application, such as a BVLOS approval, find out early what form of evidence will be useful to it.

Establishing ground truth

A trial is only as good as its reference. The usual approach is to carry an independent, high-accuracy GNSS receiver whose data is logged but never fed to the system under test. Post-processed kinematic (PPK) positioning against a nearby base station, or real-time kinematic (RTK) positioning, gives survey-grade truth under open sky, far more accurate than the system being assessed.

A few details make or break the truth data:

  • Timing. The truth receiver and the system under test need a common time base. At fixed-wing speeds, a small timing offset becomes a significant position difference.
  • Lever arms. The truth antenna and the navigation camera are not in the same place on the airframe. Measure the offset and apply it, accounting for the aircraft’s attitude.
  • Datums. Confirm that truth, reference terrain data and system output all use the same horizontal and vertical datums.
  • Truth quality. Log the truth solution’s own quality flags, and mark or exclude periods where it degraded, such as steep turns or flights near obstructions.

Simulating denial safely and legally

Radiating a jamming signal is not something to improvise. If a live-interference test is ever on the table, it needs to be planned with the relevant authorities and run under controlled conditions; check the regulatory position before going anywhere near it.

For most trials, denial is simulated in software: the GNSS input is removed from the autopilot’s navigation filter at a chosen moment while the truth receiver keeps logging. That is a legitimate test of navigation behaviour, and it allows denial to begin at precisely the point in the flight you want to examine. It is worth also testing a spoof-like case, by injecting a slowly growing offset into the GNSS input in software, to see whether the system and the autopilot detect the disagreement or follow the false position.

Log the raw camera, inertial and telemetry data from every flight. Recorded data can be replayed through later software versions, which makes comparisons between versions fair and lets you test changes without flying again.

What to measure

A single average error hides almost everything important. Useful measures include:

  • Horizontal error over time, reported as a distribution: median, a high percentile such as the 95th, and the maximum. Plot it against time since denial began.
  • Error by condition: terrain type, altitude, time of day, bank angle and phase of flight. An average across all conditions conceals exactly where the system struggles.
  • Availability: the fraction of the flight for which the system provided a usable position, and the length of the longest gap.
  • Convergence time: how long the system takes to produce a trustworthy fix from a cold start, and to recover after losing lock.
  • Confidence calibration: whether the system’s reported uncertainty matches its actual error. If it states a 95 per cent bound, errors should fall inside that bound about 95 per cent of the time.
  • Integrity: how often the system reported high confidence while its error was large. This is the most important safety measure and the easiest to overlook.
  • Closed-loop performance: once the output is fused, how well the aircraft actually tracked its flight plan, and whether transitions between sources caused visible steps in the track.
  • System health: processor load, temperature, throttling and update-rate stability, which often explain anomalies found elsewhere.

Failure cases to include

A trial that avoids the hard cases only measures the easy ones. Plan flights that deliberately include:

  • Low-feature ground such as water, uniform crops, salt pans or bare fallow, and repetitive ground such as plantation rows or dune fields.
  • Low sun, glare, haze and heat shimmer, ideally over the same route on different days.
  • Aggressive turns, climbs and descents, not just straight and level legs.
  • Altitudes at both ends of the operating band.
  • Flying close to, and past, the edge of the reference data.
  • Ground that has changed since the reference was captured.
  • Denial beginning at different points in the flight, including before take-off with no GNSS at all.
  • GNSS returning after a denial, including the simulated spoof case.
  • Mundane hardware problems: a dirty lens, heavy vibration, a hot day in a closed fuselage.

Not all of these need to pass. The point is to know how the system behaves when they happen, and in particular whether it recognises its own difficulty and reports it, rather than producing a confident wrong answer.

Running it honestly

A few habits separate a trial from a demonstration:

  • Report every flight, including aborted and inconvenient ones, and record the reason if any are excluded.
  • Repeat routes on different days, not just once.
  • Keep the raw logs, not only the summary plots.
  • Have someone who is not responsible for the system review the analysis.
  • Compare against a baseline. The same flight with an inertial-only fallback shows how much the navigation source actually adds.

Where TerrainSLAM fits

We offer tailored demonstrations of TerrainSLAM for qualified prospects, and we would rather be judged by a trial like this than by a highlight reel. TerrainSLAM provides absolute position from terrain matching, computed onboard with no satellite signal, and integrates with leading autopilot platforms without firmware changes. That means a staged evaluation, logging its output alongside GNSS and truth data before it is ever fused, can be run on the platform you already fly. If you are planning an evaluation, bring your test plan and we can work through it together.