Domain-adapted pre-trained language models for implicit information extraction in crash narratives
Researchers fine-tuned open-source pre-trained language models using Low-Rank Adaptation (LoRA) to extract structured information from free-text crash narratives in the Crash Investigation Sampling System (CISS) dataset. The method achieved 96.1% accuracy in identifying the Manner of Collision and outperformed the strongest baseline by 16.9% to 28.7% on the more complex task of extracting Crash Type.
Why it matters — It provides a highly accurate, privacy-preserving pipeline that can be deployed locally on sensitive crash data without relying on external commercial APIs, while also serving as a quality assurance tool to detect human labeling inconsistencies.
Transportation Research Part C Emerging Technologies · doi · code