TY - GEN
T1 - Schema-Driven Information Extraction from Heterogeneous Tables
AU - Bai, Fan
AU - Kang, Junmo
AU - Stanovsky, Gabriel
AU - Freitag, Dayne
AU - Dredze, Mark
AU - Ritter, Alan
N1 - Publisher Copyright:
© 2024 Association for Computational Linguistics.
PY - 2024
Y1 - 2024
N2 - In this paper, we explore the question of whether large language models can support cost-efficient information extraction from tables.We introduce schema-driven information extraction, a new task that transforms tabular data into structured records following a human-authored schema.To assess various LLM's capabilities on this task, we present a benchmark comprised of tables from four diverse domains: machine learning papers, chemistry literature, material science journals, and webpages.We use this collection of annotated tables to evaluate the ability of open-source and API-based language models to extract information from tables covering diverse domains and data formats.Our experiments demonstrate that surprisingly competitive performance can be achieved without requiring task-specific pipelines or labels, achieving F1 scores ranging from 74.2 to 96.1, while maintaining cost efficiency.Moreover, through detailed ablation studies and analyses, we investigate the factors contributing to model success and validate the practicality of distilling compact models to reduce API reliance.
AB - In this paper, we explore the question of whether large language models can support cost-efficient information extraction from tables.We introduce schema-driven information extraction, a new task that transforms tabular data into structured records following a human-authored schema.To assess various LLM's capabilities on this task, we present a benchmark comprised of tables from four diverse domains: machine learning papers, chemistry literature, material science journals, and webpages.We use this collection of annotated tables to evaluate the ability of open-source and API-based language models to extract information from tables covering diverse domains and data formats.Our experiments demonstrate that surprisingly competitive performance can be achieved without requiring task-specific pipelines or labels, achieving F1 scores ranging from 74.2 to 96.1, while maintaining cost efficiency.Moreover, through detailed ablation studies and analyses, we investigate the factors contributing to model success and validate the practicality of distilling compact models to reduce API reliance.
UR - https://www.scopus.com/pages/publications/85217621523
UR - https://www.scopus.com/pages/publications/85217621523#tab=citedBy
U2 - 10.18653/v1/2024.findings-emnlp.600
DO - 10.18653/v1/2024.findings-emnlp.600
M3 - Conference contribution
AN - SCOPUS:85217621523
T3 - EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024
SP - 10252
EP - 10273
BT - EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024
A2 - Al-Onaizan, Yaser
A2 - Bansal, Mohit
A2 - Chen, Yun-Nung
PB - Association for Computational Linguistics (ACL)
T2 - 2024 Findings of the Association for Computational Linguistics, EMNLP 2024
Y2 - 12 November 2024 through 16 November 2024
ER -