Open science

Data & Code

This is where I will share datasets, analysis code, and reproducible projects as they become available. My work uses Python and R, with explicit random seeds, pinned dependencies, and documented data sources so that results can be verified and extended.

Explicit random seeds

Reproducibility by default

Pinned dependencies

Stable environments

Documented data sources

Transparent provenance

Sensitivity analyses

Robustness checks

Available Projects

Generative AI and Everyday Life (2015-2026)

Compiled dataset of generative AI milestones, World Happiness Report life evaluation series, Gallup positive and negative experience indices, perceived freedom and stress series, perceived job security and AI attitude surveys from Gallup, Microsoft and Pew, AI exposure estimates from ILO, IMF and OECD, ChatGPT and organisational adoption series, distributional data by age, gender and geography, and experimental productivity and learning evidence used in the research brief comparing daily life before and after ChatGPT.

View report

Downloadable files

Understanding Large Language Models for Accounting and Finance (2017-2026)

Compiled dataset of LLM parameter counts and training compute by release date, LMArena (Chatbot Arena) Elo scores for frontier models, hallucination rates on document summarisation benchmarks, AI adoption rates in accounting firms from industry surveys, publisher AI policy adoption dates, global AI regulatory milestones, AI-related publication counts in top accounting journals, and Big Four AI investment data used in the research brief on how large language models work and their implications for accounting and finance research.

View report

Downloadable files

AI Tokens as an Emerging Asset Class (2020-2026)

Compiled dataset of AI token sector quarterly market capitalisation and dominance, top AI-focused blockchain projects by market cap, price history for five major AI tokens (TAO, RNDR, FET, AKT, NEAR), corporate adoption and treasury data, venture capital funding in the AI-crypto sector, comparison of US GAAP and IFRS accounting frameworks, regulatory milestone timeline, and a curated academic literature database used in the research brief on AI tokens as an emerging asset class.

View report

Downloadable files

Blockchain and Cryptocurrency in Business (2009-2026)

Compiled dataset of cryptocurrency market capitalisation and Bitcoin price history, corporate Bitcoin treasury holdings of publicly listed firms, global regulatory milestones, major exchange and protocol failure losses, spot Bitcoin and Ether ETF flow data, CBDC development progress by country, DeFi TVL and stablecoin market capitalisation trends, and a curated academic literature database used in the research brief on blockchain and cryptocurrency in business.

View report

Downloadable files

Big Data and AI in Business: NZ Tourism Nowcasting (2026)

Compiled dataset of monthly NZ international visitor arrivals (Stats NZ), Google Trends search indices for three NZ-bound travel queries, NZ tourism economic contribution from the Tourism Satellite Account, NZ AI adoption statistics, and verified academic literature on tourism demand nowcasting.

View report

Downloadable files

AI and Digital Transformation (2015-2025)

Compiled dataset of AI and digital technology adoption rates by sector and firm size, country-level digital readiness indices, corporate AI disclosure and litigation trends, AI productivity effect estimates, and regulatory development timelines used in the structured research brief on AI and digital transformation impact on business practices, firm performance, and corporate reporting.

View report

Downloadable files

Textual Analysis and NLP in Accounting and Finance

Compiled dataset of publication counts, method usage patterns, citation records, and application area trends supporting the twenty-year review of textual analysis and NLP in accounting and finance research (2005-2025).

View report

Downloadable files

What to expect

Datasets

Curated and cleaned data with documented provenance

Analysis code (Python / R)

Reproducible scripts with pinned environments

Replication materials

Full replication packages with explicit random seeds

Methodological notes

Detailed documentation of design choices and sensitivity analyses

Read reflections on this work

Read insights