Shubham Tibrewal

Independent Researcher · Chartered Accountant

LLM Evaluation Multilingual Reasoning Low-Resource Languages Reproducibility
Tools & Software
Eval Stats Toolkit
Wilson CIs, Newcombe difference-of-proportions, and McNemar's paired test for LLM evaluation results.
doi:10.5281/zenodo.21268412
Open tool →
LLM Memory Calculator
Memory and VRAM requirements for language models at every quantization level — FP32 to 1.58-bit.
doi:10.5281/zenodo.21268419
Open tool →
McNemar's Test for NLP
Paired significance test for comparing two models on the same test set. Exact binomial fallback when b+c < 25.
doi:10.5281/zenodo.21268428
Open tool →
Research in Progress
Cross-Lingual Reasoning Gap: Math Problem-Solving Across English, Hindi, Tamil, and Sanskrit Under review
Shubham Tibrewal · 2026 · Targeting MRL @ EMNLP 2026
Measures GSM8K math-reasoning accuracy degradation when problems are translated from English into Hindi, Tamil, and Sanskrit, comparing direct prompting against chain-of-thought. Includes manual Sanskrit verification — a documented confound in automated translation pipelines for classical scripts.
About

I am an independent AI researcher and Chartered Accountant, building a research identity focused on LLM evaluation methodology, multilingual reasoning, and reproducibility in NLP. My work sits at the intersection of rigorous statistical evaluation and low-resource language coverage — areas where commercial incentives leave systematic gaps.

All tools are open-source, single-file, dependency-free, and run entirely in the browser. Code and datasets are published on GitHub and archived on Zenodo with DOIs for citation.