Abstract
Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by. Existing table extraction datasets often focus on scientific tables due to the vast amount of academic articles that are readily available, along with their source code. However, there are significant layout and typographical differences between tables found across scientific, financial, and other domains. Current datasets often lack the words, and their positions, contained within the tables, instead relying on unreliable OCR to extract these features for training modern machine learning models on natural language processing tasks. Therefore, there is a need for a more general method of obtaining labelled data. We present SynFinTabs, a large-scale, labelled dataset of synthetic financial tables. Our hope is that our method of generating these synthetic tables is transferable to other domains. To demonstrate the effectiveness of our dataset in training models to extract information from table images, we create FinTabQA, a layout large language model trained on an extractive question-answering task. We test our model using real-world financial tables and compare it to a state-of-the-art generative model and discuss the results. We make the dataset, model, and dataset generation code publicly available (https://ethanbradley.co.uk/research/synfintabs).
| Original language | English |
|---|---|
| Title of host publication | Document Analysis and Recognition – ICDAR 2025 Workshops: Proceedings, Part II |
| Editors | Lianwen Jin, Richard Zanibbi, Veronique Eglin |
| Place of Publication | Cham |
| Publisher | Springer Nature Switzerland |
| Chapter | 6 |
| Pages | 85–100 |
| Number of pages | 16 |
| Volume | 2 |
| ISBN (Electronic) | 9783032093714 |
| ISBN (Print) | 9783032093707 |
| DOIs | |
| Publication status | Published - 02 Jan 2026 |
| Event | International Workshops co-located with the 19th International Conference on Document Analysis and Recognition, ICDAR 2025 - Wuhan, China Duration: 20 Sept 2025 → 21 Sept 2025 |
Publication series
| Name | Lecture Notes in Computer Science |
|---|---|
| Publisher | Springer |
| Number | 1 |
| Volume | 16226 |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | International Workshops co-located with the 19th International Conference on Document Analysis and Recognition, ICDAR 2025 |
|---|---|
| Country/Territory | China |
| City | Wuhan |
| Period | 20/09/2025 → 21/09/2025 |
Bibliographical note
8 figuresKeywords
- Synthetic data
- Information extraction
- Table extraction
ASJC Scopus subject areas
- Theoretical Computer Science
- General Computer Science
Fingerprint
Dive into the research topics of 'SynFinTabs: a dataset of synthetic financial tables for information and table extraction'. Together they form a unique fingerprint.Datasets
-
SynFinTabs
Bradley, E. (Creator), Roman, M. (Creator), Rafferty, K. (Creator) & Devereux, B. (Creator), Hugging Face, 05 Dec 2024
DOI: 10.57967/hf/3741, https://huggingface.co/datasets/ethanbradley/synfintabs
Dataset
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver