Projects
Everything I build in public, grouped by what it is for.
From AI evaluation and agent controls to clinical genomics data tooling and applied machine learning. Every link goes to the code on GitHub.
Building in the open
Three connected projects, one question: how do you get real value from AI without losing control of it?
-
Act safely
agent-guardrails
How do you let an AI agent act without it sending the wrong thing, twice, to the wrong people?
Controls for AI agents that write or send: allow-listed actions, draft and approve modes, execution-time re-validation, idempotency, a tamper-evident audit log and a kill switch.
-
Measure quality
evidence-synthesis-eval
Does an AI summary of the evidence say what the sources say? What did it miss, and what did checking it cost?
Evaluation harness for AI-assisted evidence synthesis: faithfulness, coverage, misstatements and human checking effort. Built on Inspect from the UK AI Security Institute.
-
Assure and decide
ai-assurance-kit
How can a team do proportionate AI assurance in an afternoon rather than a quarter?
Turns an AI use case into a validated, reviewable record and a one-page card, with checks drawn from public UK guidance and a NIST AI RMF crosswalk.
Clinical genomics and biobank data
Sanitised, generalised versions of tooling from my NIHR BioResource work, rebuilt as tested Python packages. Synthetic data only.
HLA and genotyping
-
hla-pipeline-manager
Orchestrates HLA imputation at 50,000+ sample scale: batch execution, verification, safe deployment, clinical reports.
-
hla-imputation-analyst
Quality control and forensic analysis of HLA imputation outputs.
-
hla-variant-investigator
Carrier identification, dosage auditing and imputation quality checks for HLA datasets.
-
apoe-genotyping-toolkit
APOE genotype calling, trial feasibility estimates and stratified recall lists.
Data preparation and QC
-
gwas-data-preparation
Multi-batch genotype merging, standard QC filters and strand resolution for GWAS.
-
genomic-qc-toolkit
Sample and variant QC for WGS, WES and array data, with cross-platform concordance.
-
vcf-plink-converter
Validated VCF and PLINK conversion wrapping PLINK and BCFtools.
-
ld-linkage-mapper
Finds linkage-disequilibrium proxies when a target variant is not on the array.
-
snp-feasibility-checker
Checks SNP coverage across arrays and estimates recall-study yield.
Cohorts and recall
-
biobank-variant-explorer
Audits variant presence across genotyping arrays by scanning PLINK files.
-
clinical-cohort-selector
Genotype-stratified, demographically balanced cohort selection for recall studies.
-
recall-study-generator
Designs recall-by-genotype studies with eligibility filtering and protocol generation.
Governance and secure delivery
-
genomic-cohort-delivery-pipeline
Assembles, corrects and delivers multi-batch genomic cohorts.
-
biobank-data-release-manager
Controlled data releases with VCF extraction and sample concordance checks.
-
secure-genomic-transfer
GPG-encrypted transfer with SHA-256 manifests and an audit trail.
-
biobank-sar-toolkit
GDPR subject access requests: participant discovery, ID mapping and reporting.
Clinical AI platform
-
gut-reaction-platform
Reference architecture for an inflammatory bowel disease research platform: rule-based clinical NLP phenotyping, a PII auditor with a mocked model call, data harmonisation services and a dashboard, with Kubernetes and Terraform definitions.
Applied machine learning
-
fraud-detection-system
Value-weighted fraud detection with CatBoost, isotonic calibration and expected-value ranking.
-
retail-demand-forecasting-at-scale
Demand forecasting pipeline with gradient boosting, feature engineering and backtesting.
-
british-invoice-digitization
Computer vision for invoice field extraction with YOLOv5 and a FastAPI service.
Standards across these repositories
- Each project is a tested Python package with continuous integration, and a
docs/WHY.mdthat explains its design decisions. - Original shell scripts sit in
legacy/beside the rewrites, to show how ad-hoc tooling became maintainable software. - No real participant data, credentials or proprietary information. Synthetic data only.
- I build with AI coding assistants and say so: commits written with Claude carry a co-author line. I review, test and own everything that is merged.