Fig. 1 Local inference unit

Small models,
local machines. Making a one-billion-parameter model trustworthy enough to read a contract.

I'm Ryan — an accountant turned broker turned AI researcher. I work on the gap between a model that answers and a model you'd stake a deal on.

DegreeMSc AI & Data Science, Distinction
FocusQuantized SLMs · RAG · Reliability
BasedHull, England
1 — AXIAL FAN 2 — 8-PIN POWER 3 — PCIe ×16
Fig. 1 · Local inference unit 0000 RPM
§ 01 The short version

I spent nine years reading contracts before I ever tried to teach a model to.

Accounting first — SOX documentation, payroll, financial models. Then eight years as a broker-associate, where the real job turned out to be reading paper carefully and quickly, and where I built my first OCR pipeline because I was tired of retyping the same twelve fields.

That's the thread. I know what it costs when a document is read wrong, so I'm sceptical of systems that are fluent but unaccountable. My MSc research asked how far you can push a quantized 1B-parameter model with retrieval, fine-tuning and output validation before it becomes safe enough for legal, accounting and real-estate review.

The answer, roughly: much further than expected, and never far enough to remove the human.

A model that runs on your own machine keeps the contract on your own machine. That property is worth engineering around.
Fig. 2 · Stage I

The model drafts.

A quantized Llama 3.2 1B, running locally, extracts and assembles the clause structure. Fast, cheap, private — and, unsupervised, confidently wrong often enough to matter.

Fig. 2 · Stage II

Retrieval grounds it.

Four retrieval strategies compared against the CUAD corpus. Every generated span has to point back to text that actually exists in the document.

Fig. 2 · Stage III

Counsel marks it up.

Validation is the red pen: schema checks, span verification, an LLM-as-judge pass. Anything that can't be traced to the source gets flagged, not published.

Fig. 2 · Result

+68% accuracy over baseline.

A 2³ factorial design across RAG, QLoRA fine-tuning and output validation. Statistically significant by Wilcoxon signed-rank, p = 0.001. The human still signs.

? § REVIEWED HUMAN COUNSEL · FINAL
§ 02 Record of work

Eleven years of reading documents for a living, then teaching machines to.

May 2026 — present

Machine Learning Technician

Independent research engagement with Dr. Eva J. Muir · Remote
  • Building an automated bioacoustic detection pipeline for moon bear and sun bear vocalizations across thousands of hours of passive acoustic monitoring audio from the Laotian jungle — against a dense background of overlapping non-target species.
  • Engineered the feature extraction stack (MFCCs, mel-spectrogram statistics, spectral centroid, zero-crossing rate) and trained a class-weighted Random Forest to flag candidates; CNN classification in development.
  • Built a sliding-window scanner over the Google Drive API to process large archives and export a confidence-ranked detection log for expert review.
May 2025 — June 2026

AI Graduate Student Researcher

University of Hull · Kingston-upon-Hull, England
  • +68% accuracyp = 0.001 Designed a 2³ factorial experiment across RAG, QLoRA fine-tuning and output validation on a quantized Llama 3.2 1B against the CUAD contracts dataset; the best treatment beat baseline significantly on Wilcoxon signed-rank.
  • Compared four RAG retrieval strategies and QLoRA-tuned the model on a rented A100, with results validated through an LLM-as-a-judge evaluation pipeline.
Feb 2018 — May 2025

Broker-Associate & Technology Lead

Keller Williams Realty, Inc. · Tampa, FL
  • Implemented the team's first CRM, replacing ad-hoc spreadsheet tracking with a central platform for client and transaction management.
  • Built a Python pipeline using pyTesseract OCR and NLP text extraction to pull key terms from contracts on local machines and auto-generate templated client emails — the earliest version of the problem I now research.
  • Developed a proprietary data-driven pricing strategy and mentored six team members on analytics-based valuation.
Jul 2017 — Jan 2018

Accountant

C&L Value Advisors LLC · Tampa, FL
  • Prepared payroll and financial statements for a portfolio of small and mid-sized business clients in QuickBooks and Thomson Reuters; built Excel models for company-wide forecasting and scenario analysis.
Oct 2016 — May 2017

Internal Controls Intern

Ditech Financial LLC · Tampa, FL
  • Documented internal controls for SOX audit compliance and developed process flowcharts for risk assessment; worked with external auditors on annual control opinions and remediation.
  • 150 hrs → 30 min Built a VBA tool in Excel to consolidate accounts ahead of journal entries in SAP.
§ 03 Contests & honours

Where the work has been put in front of judges.

2026

Distinction — MSc Artificial Intelligence & Data Science

University of Hull

The highest classification in the UK system, roughly equivalent to summa cum laude.

Distinction
2026

Invited speaker — UK Real Estate Investment & Infrastructure Forum

UKREiiF · Leeds, England

Presented AI research on multi-agent contract review to a property and infrastructure audience.

Invited
2025

IBM AI Design Challenge

Eight competing teams · October 2025

Took the challenge outright and was named Most Valuable Player of the competition.

Winner · MVP
2025

AVEVA EcoTech Emerge

Net Zero innovation challenge

Finished in the top twenty of the field.

Top 20 finalist
2023–24

Top Teams

Tampa Bay Business Journal

Named across two consecutive years of the journal's regional ranking.

Listed ×2
2012

University of Florida Programming Contest (Java)

High school programming team

Second place — the first time any of this was competitive.

2nd place
§ 04 Instruments

Languages

Python · SQL · R · JavaScript · VBA

Data science & ML

Pandas · NumPy · Scikit-learn · TensorFlow · Random Forest · CNNs

NLP & LLM tooling

Llama 3.2 · Mistral · QLoRA / LoRA fine-tuning · Hugging Face Transformers · RAG pipeline design

Specialised

Tesseract OCR · ChromaDB · Label Studio · Statistical modelling · Google Data Analytics Certificate

§ 05 Writing

Longer pieces, on whatever currently has my attention.

A CV is a narrow instrument. This is the wider one — essays, arguments I'm still working out, and things with nothing to do with any of it.

§ 06 Verse & personal writing

Poems, and the writing that isn't for anything.

Line breaks kept exactly as written. Nothing here is trying to be useful.

§ 07 Marginalia

Shorter thoughts, filed as they arrive.

Observations, half-arguments, things overheard. Too small for an essay, too good to lose.

§ 08 Correspondence

Get in touch.

Open to research collaboration, applied work on small-model reliability, and most conversations. I answer everything that isn't automated.