Product definition
The short product is issued at eight UTC cycles and is valid for nine hours. The long product is issued at four UTC cycles and is valid for 30 hours. Records with malformed validity windows are excluded.
The current model predicts prevailing wind, gust, visibility, weather, cloud cover and cloud base. It does not yet predict reliable onset and end times for TEMPO or BECMG groups, so those groups are not generated.
Data acquisition and parsing
Historical observations come from VABB METAR archives collected from Iowa Environmental Mesonet and Ogimet. Issued TAFs were extracted from Ogimet, parsed into forecast elements, and retained with raw text, issue time and validity metadata. Unsupported validity periods and parsing failures are excluded.
The current inventory contains 482,141 observation rows and 55,823 issued TAF texts through 2 August 2026. Rare events remain limited; thunderstorm results must therefore be interpreted separately from common conditions.
Issue-time lineage
Every inference row is keyed by station, product and issue time. The latest METAR must have a receipt time no later than the selected issue time, and the previous TAF must have been issued earlier. The backend repeats these checks for live cycles and refuses to invent data when a source is unavailable.
Open-Meteo historical weather or reanalysis is not treated as an operational forecast archive. Raw archived IMD NWP runs with initialization and availability timestamps remain a deployment gate.
Model and TAF construction
Separate gradient-boosting models predict continuous, binary and categorical forecast elements. Thresholds and any combination with the previous issued TAF are selected using the 2023 validation period. Results from 2024 onward are not used to choose a prediction for an individual case.
A rule-based compiler converts predicted elements into TAF text and checks wind formatting, weather-visibility consistency, cloud syntax, and parser round trips.
Train, validation and test
The structured score averages 14 atom-level components including visibility and cloud categories, binary weather/change flags, wind direction within 30 degrees and wind speed within 3 knots. The generated text is separately measured with token F1 and deterministic syntax validation.
The corrected archive retains co-issued short and long products and partitions previous-TAF features by product type. This correction reduced the combined score below 80%, so the target is not achieved. Token F1 is shown in the dashboard only as lexical similarity to the issued TAF. Proper acceptance also requires subsequent METAR/SPECI verification, hazard recall and false-alarm analysis, calibration, seasonal breakdowns and prospective shadow evaluation.
Known limitations
- No raw archived IMD NWP run-level fields are ingested.
- Thunderstorm model recall is zero on the frozen test; persistence supplies the current ensemble atom.
- No validated timing heads exist for TEMPO, BECMG or probability groups.
- Historical human TAF text is a forecast reference, not observed weather truth.
- The test period has been inspected repeatedly; future model selection needs rolling-origin or prospective evaluation.
- Human forecaster review and shadow-mode evidence are required before any operational recommendation.
Reference sources
- NOAA Localized Aviation MOS ProgramOperational reference for station-level aviation guidance.
- NOAA LAMP verification FAQElement and category verification methods.
- Met Office TAF verificationVerification that accounts for rare categories.
- Open-Meteo Single Runs APIArchived forecast reference used only for prototype work.
- Aviation Weather Center Data APIRecent TAF and METAR comparison feed.
- IMD NWP products sitemapPublic reference for IMD forecast products.