← Back to full report

Textual Analysis and Natural Language Processing in Accounting and Finance Research: A Twenty-Year Review (2005-2025)

One-Page Summary

Dr Yuqian Zhang · July 2026

What This Report Is About

This review traces the development of textual analysis and natural language processing (NLP) in accounting and finance research over two decades. It covers the data sources researchers use, the methods they deploy, the applications they pursue, and the validation challenges they face. The review maps a field that has grown from about 2 papers per year in 2005 to approximately 90 in 2023.

The narrative follows six methodological phases: general-purpose dictionaries, domain-specific word lists, machine learning classifiers, topic models, word embeddings, transformer models (FinBERT), and large language models (GPT). Each phase expanded the kinds of questions researchers could ask of corporate language.

Key Findings

The Loughran-McDonald (2011) paper changed the field. With approximately 5,700 Google Scholar citations, it is the most-cited paper in textual analysis research. Its central insight was that general-purpose dictionaries systematically misclassify financial language: words such as "liability," "tax," and "cost" are coded as negative in general dictionaries but are neutral in financial contexts. The LM Master Dictionary became standard research infrastructure for the field.
The performance gap between dictionaries and transformers is large. On sentiment classification, dictionary methods achieve about 62 per cent accuracy, while domain-trained transformers such as FinBERT reach about 88 per cent. This 26-percentage-point gap is large enough to change the conclusions of applied studies. The fastest growth in the field occurred after 2017, as machine learning methods entered mainstream accounting research.
Textual signals carry information beyond the numbers. Disclosure tone predicts future returns and operating performance. Negative tone consistently has stronger and more persistent effects than positive tone, consistent with the greater credibility of unfavourable information in disclosure settings. Less readable 10-K filings are associated with lower earnings persistence.

Key Statistics

From ~2 papers/year (2005) to ~90 papers/year (2023) in top journals
61% of papers published in accounting journals; 31% in finance journals
~5,700 citations for Loughran and McDonald (2011)
37% of all textual analysis papers involve tone or sentiment measurement
88.2%: FinBERT sentiment classification accuracy vs. 62.1% for dictionaries
6 phases of NLP method progression: dictionaries through to LLMs (2023-present)
6 word lists in the LM Master Dictionary: negative, positive, uncertainty, litigious, strong modal, weak modal
96% accuracy: generative LLMs on non-answer detection in earnings calls

Methodological Concerns

As methods grow more complex, validation becomes more important. Key concerns include measurement validity (does the NLP measure capture the construct it claims to capture?), out-of-sample testing (does the model work on data it has not seen?), reproducibility (can other researchers replicate the result?), and look-ahead bias (are models inadvertently trained on data from the future, particularly a concern for LLMs pre-trained on web-scale corpora?). The review argues for systematic human coding benchmarks as a validation standard for new NLP measures in accounting research.

Why It Matters

Textual analysis has moved from the periphery to the centre of empirical accounting and finance research. The questions it addresses (what do firms say? how do they say it? what does that language tell us about future outcomes?) are fundamental to understanding disclosure, corporate communication, and capital markets. For doctoral students and early-career researchers, understanding this methodological trajectory is essential, as the tools available today are qualitatively different from those available a decade ago.