This study generates short 9-hour and long 30-hour TAFOR guidance for VABB using weather information available at issue time. The present system is an evaluated research prototype, not an operational aviation product.

78.983%combined structured component score
80.148%short 9h
77.288%long 30h
30.10%generated-text token F1

The structured figures compare selected forecast atoms with historical human TAF atoms. They are not operational accuracy. Short 9h exceeds 80% by only 0.008 percentage points, and the inspected post-2024 period is not an independent confirmation set.

01

Product definition

The short product is issued at eight UTC cycles and is valid for nine hours. The long product is issued at four UTC cycles and is valid for 30 hours. Records with malformed validity windows are excluded.

Output scope

The current model predicts prevailing wind, gust, visibility, weather, cloud cover and cloud base. It does not yet predict reliable onset and end times for TEMPO or BECMG groups, so those groups are not generated.

02

Data acquisition and parsing

Historical observations come from VABB METAR archives collected from Iowa Environmental Mesonet and Ogimet. Issued TAFs were extracted from Ogimet, parsed into forecast elements, and retained with raw text, issue time and validity metadata. Unsupported validity periods and parsing failures are excluded.

The current inventory contains 482,141 observation rows and 55,823 issued TAF texts through 2 August 2026. Rare events remain limited; thunderstorm results must therefore be interpreted separately from common conditions.

03

Issue-time lineage

Every inference row is keyed by station, product and issue time. The latest METAR must have a receipt time no later than the selected issue time, and the previous TAF must have been issued earlier. The backend repeats these checks for live cycles and refuses to invent data when a source is unavailable.

METAR historyreceived by issue time
Feature builder61 allowed fields
Atom modelsshort / long suites
TAF compilergrammar constrained

Open-Meteo historical weather or reanalysis is not treated as an operational forecast archive. Raw archived IMD NWP runs with initialization and availability timestamps remain a deployment gate.

04

Model and TAF construction

Separate gradient-boosting models predict continuous, binary and categorical forecast elements. Thresholds and any combination with the previous issued TAF are selected using the 2023 validation period. Results from 2024 onward are not used to choose a prediction for an individual case.

A rule-based compiler converts predicted elements into TAF text and checks wind formatting, weather-visibility consistency, cloud syntax, and parser round trips.

05

Train, validation and test

Through 202233,802 training rows
20233,941 validation rows
2024 onward8,412 test rows

The structured score averages 14 atom-level components including visibility and cloud categories, binary weather/change flags, wind direction within 30 degrees and wind speed within 3 knots. The generated text is separately measured with token F1 and deterministic syntax validation.

The corrected archive retains co-issued short and long products and partitions previous-TAF features by product type. This correction reduced the combined score below 80%, so the target is not achieved. Token F1 is shown in the dashboard only as lexical similarity to the issued TAF. Proper acceptance also requires subsequent METAR/SPECI verification, hazard recall and false-alarm analysis, calibration, seasonal breakdowns and prospective shadow evaluation.

06

Known limitations

  • No raw archived IMD NWP run-level fields are ingested.
  • Thunderstorm model recall is zero on the frozen test; persistence supplies the current ensemble atom.
  • No validated timing heads exist for TEMPO, BECMG or probability groups.
  • Historical human TAF text is a forecast reference, not observed weather truth.
  • The test period has been inspected repeatedly; future model selection needs rolling-origin or prospective evaluation.
  • Human forecaster review and shadow-mode evidence are required before any operational recommendation.
07

Reference sources