Text Classification
Transformers
Safetensors
English
bert
token-classification
financial NLP
named entity recognition
sequence labeling
structured extraction
hierarchical taxonomy
XBRL
iXBRL
SEC filings
financial-information-extraction
text-embeddings-inference
Instructions to use AAU-NLP/BERT-SL1000 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AAU-NLP/BERT-SL1000 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AAU-NLP/BERT-SL1000")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("AAU-NLP/BERT-SL1000") model = AutoModelForTokenClassification.from_pretrained("AAU-NLP/BERT-SL1000", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from AAU-NLP/BERT-SL1000: direct link, hf CLI and curl.
- Browser
- Download file 4.34 kB
-
https://huggingface.co/AAU-NLP/BERT-SL1000/resolve/main/README.md
- Command line
-
hf download hf://AAU-NLP/BERT-SL1000/README.md
-
curl -L -o README.md https://huggingface.co/AAU-NLP/BERT-SL1000/resolve/main/README.md
4.34 kB
metadata
base_model: bert-base-uncased
datasets:
- AAU-NLP/HiFi-KPI
language:
- en
library_name: transformers
license: apache-2.0
model_name: BERT-SL1000
pipeline_tag: text-classification
tags:
- financial NLP
- named entity recognition
- sequence labeling
- structured extraction
- hierarchical taxonomy
- XBRL
- iXBRL
- SEC filings
- financial-information-extraction
task_categories:
- text-classification
- token-classification
task_ids:
- named-entity-recognition
- financial-information-extraction
pretty_name: 'BERT-SL1000: Sequence Labeling for Financial KPI Extraction'
size_categories: 1M<n<10M
languages:
- en
dataset_name: HiFi-KPI
model_description: >
BERT-SL1000 is a **BERT-based sequence labeling model** fine-tuned on the
**HiFi-KPI dataset** for extracting
**financial key performance indicators (KPIs)** from **SEC earnings filings
(10-K & 10-Q)**. It specializes in identifying
entities, such as revenue, earnings, and financial ratios, using **token
classification**.
This model is part of the **HiFi-KPI benchmark** and is optimized for
**hierarchical label consistency**.
dataset_link: https://huggingface.co/datasets/AAU-NLP/HiFi-KPI
repo_link: https://github.com/aaunlp/HiFi-KPI
BERT-SL1000
Model Description
BERT-SL1000 is a BERT-based sequence labeling model fine-tuned on the HiFi-KPI dataset for extracting financial key performance indicators (KPIs) from SEC earnings filings (10-K & 10-Q). It specializes in identifying entities, such as revenue, earnings etc.
This model was introduced in the paper HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings by Rasmus Aavang, Giovanni Rizzi, Rasmus Bøggild, Alexandre Iolov, Mike Zhang, and Johannes Bjerva.
Use Cases
- Extracting financial KPIs from SEC 10-K and 10-Q reports
- Financial document parsing with iXBRL-based entity recognition
Performance
- Trained on 1,000 most frequent labels from the HiFi-KPI dataset
Dataset & Code
- Dataset: HiFi-KPI on Hugging Face
- Code: HiFi-KPI GitHub Repository
Citation
@inproceedings{aavang-etal-2026-hifi,
title = "{H}i{F}i-{KPI}: A Dataset for Hierarchical {KPI} Extraction from Earnings Filings",
author = "Aavang, Rasmus T. and
Rizzi, Giovanni and
Tjalk-B{\o}ggild, Rasmus and
Iolov, Alexandre and
Zhang, Mike and
Bjerva, Johannes",
editor = "Piperidis, Stelios and
Bel, N{\'u}ria and
van den Heuvel, Henk and
Ide, Nancy and
Krek, Simon and
Toral, Antonio",
booktitle = "Proceedings of the Fifteenth Language Resources and Evaluation Conference",
month = may,
year = "2026",
address = "Palma de Mallorca, Spain",
publisher = "ELRA Language Resource Association",
url = "https://aclanthology.org/2026.lrec-1.30/",
doi = "10.63317/2nbsp7zzfb3g",
pages = "441--455",
abstract = "Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 2.5K-instance subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured extraction. Finally, a qualitative analysis reveals that extraction errors primarily relate to dates. We open-source all code and data at Anonymous."
}