I'm Ryan — an accountant turned broker turned AI researcher. I work on the gap between a model that answers and a model you'd stake a deal on.
Accounting first — SOX documentation, payroll, financial models. Then eight years as a broker-associate, where the real job turned out to be reading paper carefully and quickly, and where I built my first OCR pipeline because I was tired of retyping the same twelve fields.
That's the thread. I know what it costs when a document is read wrong, so I'm sceptical of systems that are fluent but unaccountable. My MSc research asked how far you can push a quantized 1B-parameter model with retrieval, fine-tuning and output validation before it becomes safe enough for legal, accounting and real-estate review.
The answer, roughly: much further than expected, and never far enough to remove the human.
A quantized Llama 3.2 1B, running locally, extracts and assembles the clause structure. Fast, cheap, private — and, unsupervised, confidently wrong often enough to matter.
Four retrieval strategies compared against the CUAD corpus. Every generated span has to point back to text that actually exists in the document.
Validation is the red pen: schema checks, span verification, an LLM-as-judge pass. Anything that can't be traced to the source gets flagged, not published.
A 2³ factorial design across RAG, QLoRA fine-tuning and output validation. Statistically significant by Wilcoxon signed-rank, p = 0.001. The human still signs.
The highest classification in the UK system, roughly equivalent to summa cum laude.
Presented AI research on multi-agent contract review to a property and infrastructure audience.
Took the challenge outright and was named Most Valuable Player of the competition.
Finished in the top twenty of the field.
Named across two consecutive years of the journal's regional ranking.
Second place — the first time any of this was competitive.
Python · SQL · R · JavaScript · VBA
Pandas · NumPy · Scikit-learn · TensorFlow · Random Forest · CNNs
Llama 3.2 · Mistral · QLoRA / LoRA fine-tuning · Hugging Face Transformers · RAG pipeline design
Tesseract OCR · ChromaDB · Label Studio · Statistical modelling · Google Data Analytics Certificate
A CV is a narrow instrument. This is the wider one — essays, arguments I'm still working out, and things with nothing to do with any of it.
Line breaks kept exactly as written. Nothing here is trying to be useful.
Observations, half-arguments, things overheard. Too small for an essay, too good to lose.