Leaderboard · Plutus

Open Greek Financial LLM Leaderboard

Greek is a low-resource language for finance. We test general, Greek-adapted and financial LLMs on five Greek financial tasks from the Plutus benchmark, from reading annual reports to classifying news headlines.

0models
0tasks
0best average
Top 8 · average scoreFull ranking →
Findings

What the numbers say

Ranking

Who leads, task by task

Switch the metric to re-rank. Hover a bar for every score. Click a model family to hide it.

Size

Does a bigger model mean better Greek finance?

Open models only, by parameter count (log scale). API models are left out because their sizes are not public.

Spread

Where models struggle

Each dot is one model on one task. Hover or tap a dot to trace that model across all tasks.

Matrix

Every score at a glance

Data

All results

Scores are on a 0–100 scale; for classification and QA they are normalized so that random guessing scores 0. Raw scores are in the CSV.

People & partners

Built together

Collaborators

Author affiliations on the Plutus paper.

Acknowledgments

Data and annotation work this benchmark builds on.

Partners NaCTeM AIRC

Method

Five Greek financial tasks from Plutus-ben, run with the FinBen evaluation framework. Average is the unweighted mean of the five task scores.

Add a model

Open an issue on The-FinAI/FinBen with the model's Hub id and revision. We run the same evaluation and add the scores here.

Cite

@misc{peng2025plutusbenchmarkinglargelanguage,
  title={Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance},
  author={Xueqing Peng and Triantafillos Papadopoulos and Efstathia Soufleri and Polydoros Giannouris and Ruoyu Xiang and Yan Wang and Lingfei Qian and Jimin Huang and Qianqian Xie and Sophia Ananiadou},
  year={2025}, eprint={2502.18772}, archivePrefix={arXiv}, primaryClass={cs.CL},
  url={https://arxiv.org/abs/2502.18772}
}