Open Greek Financial LLM Leaderboard
Greek is a low-resource language for finance. We test general, Greek-adapted and financial LLMs on five Greek financial tasks from the Plutus benchmark, from reading annual reports to classifying news headlines.
What the numbers say
Who leads, task by task
Switch the metric to re-rank. Hover a bar for every score. Click a model family to hide it.
Does a bigger model mean better Greek finance?
Open models only, by parameter count (log scale). API models are left out because their sizes are not public.
Where models struggle
Each dot is one model on one task. Hover or tap a dot to trace that model across all tasks.
Every score at a glance
All results
Scores are on a 0–100 scale; for classification and QA they are normalized so that random guessing scores 0. Raw scores are in the CSV.
Built together
Collaborators
Author affiliations on the Plutus paper.
Acknowledgments
Data and annotation work this benchmark builds on.
Method
Five Greek financial tasks from Plutus-ben, run with the FinBen evaluation framework. Average is the unweighted mean of the five task scores.
Add a model
Open an issue on The-FinAI/FinBen with the model's Hub id and revision. We run the same evaluation and add the scores here.
Cite
@misc{peng2025plutusbenchmarkinglargelanguage,
title={Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance},
author={Xueqing Peng and Triantafillos Papadopoulos and Efstathia Soufleri and Polydoros Giannouris and Ruoyu Xiang and Yan Wang and Lingfei Qian and Jimin Huang and Qianqian Xie and Sophia Ananiadou},
year={2025}, eprint={2502.18772}, archivePrefix={arXiv}, primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.18772}
}
The Fin AI



