Projects

Everything I build in public, grouped by what it is for.

From AI evaluation and agent controls to clinical genomics data tooling and applied machine learning. Every link goes to the code on GitHub.

Building in the open

Three connected projects, one question: how do you get real value from AI without losing control of it?

  • Act safely

    agent-guardrails

    How do you let an AI agent act without it sending the wrong thing, twice, to the wrong people?

    Controls for AI agents that write or send: allow-listed actions, draft and approve modes, execution-time re-validation, idempotency, a tamper-evident audit log and a kill switch.

    • Python
    • SQLite
    • Audit log
    • Claude tool use
  • Measure quality

    evidence-synthesis-eval

    Does an AI summary of the evidence say what the sources say? What did it miss, and what did checking it cost?

    Evaluation harness for AI-assisted evidence synthesis: faithfulness, coverage, misstatements and human checking effort. Built on Inspect from the UK AI Security Institute.

    • Python
    • Inspect
    • Model-graded scorers
    • Deterministic scorers
  • Assure and decide

    ai-assurance-kit

    How can a team do proportionate AI assurance in an afternoon rather than a quarter?

    Turns an AI use case into a validated, reviewable record and a one-page card, with checks drawn from public UK guidance and a NIST AI RMF crosswalk.

    • Python
    • Pydantic
    • JSON Schema
    • Jinja2

Clinical genomics and biobank data

Sanitised, generalised versions of tooling from my NIHR BioResource work, rebuilt as tested Python packages. Synthetic data only.

HLA and genotyping

  • hla-pipeline-manager

    Orchestrates HLA imputation at 50,000+ sample scale: batch execution, verification, safe deployment, clinical reports.

    • HLA imputation
    • Batch orchestration
    • Clinical reports
  • hla-imputation-analyst

    Quality control and forensic analysis of HLA imputation outputs.

    • HLA imputation
    • Quality control
  • hla-variant-investigator

    Carrier identification, dosage auditing and imputation quality checks for HLA datasets.

    • HLA
    • Carrier identification
    • Dosage audit
  • apoe-genotyping-toolkit

    APOE genotype calling, trial feasibility estimates and stratified recall lists.

    • APOE
    • Genotype calling
    • Recall lists

Data preparation and QC

  • gwas-data-preparation

    Multi-batch genotype merging, standard QC filters and strand resolution for GWAS.

    • GWAS
    • Genotype QC
    • Strand resolution
  • genomic-qc-toolkit

    Sample and variant QC for WGS, WES and array data, with cross-platform concordance.

    • WGS
    • WES
    • Arrays
    • Concordance
  • vcf-plink-converter

    Validated VCF and PLINK conversion wrapping PLINK and BCFtools.

    • VCF
    • PLINK
    • BCFtools
  • ld-linkage-mapper

    Finds linkage-disequilibrium proxies when a target variant is not on the array.

    • Linkage disequilibrium
    • Proxy variants
  • snp-feasibility-checker

    Checks SNP coverage across arrays and estimates recall-study yield.

    • SNP arrays
    • Coverage
    • Recall feasibility

Cohorts and recall

  • biobank-variant-explorer

    Audits variant presence across genotyping arrays by scanning PLINK files.

    • Variant audit
    • Genotyping arrays
    • PLINK
  • clinical-cohort-selector

    Genotype-stratified, demographically balanced cohort selection for recall studies.

    • Cohort selection
    • Stratification
    • Recall studies
  • recall-study-generator

    Designs recall-by-genotype studies with eligibility filtering and protocol generation.

    • Recall-by-genotype
    • Eligibility
    • Study design

Governance and secure delivery

  • genomic-cohort-delivery-pipeline

    Assembles, corrects and delivers multi-batch genomic cohorts.

    • Cohort delivery
    • Multi-batch
  • biobank-data-release-manager

    Controlled data releases with VCF extraction and sample concordance checks.

    • Data release
    • VCF
    • Concordance
  • secure-genomic-transfer

    GPG-encrypted transfer with SHA-256 manifests and an audit trail.

    • GPG
    • SHA-256
    • Audit trail
  • biobank-sar-toolkit

    GDPR subject access requests: participant discovery, ID mapping and reporting.

    • GDPR
    • Subject access requests
    • ID mapping

Clinical AI platform

  • gut-reaction-platform

    Reference architecture for an inflammatory bowel disease research platform: rule-based clinical NLP phenotyping, a PII auditor with a mocked model call, data harmonisation services and a dashboard, with Kubernetes and Terraform definitions.

    • Clinical NLP
    • OMOP
    • React
    • Kubernetes
    • Terraform

Applied machine learning

  • fraud-detection-system

    Value-weighted fraud detection with CatBoost, isotonic calibration and expected-value ranking.

    • CatBoost
    • Calibration
    • Expected value
  • retail-demand-forecasting-at-scale

    Demand forecasting pipeline with gradient boosting, feature engineering and backtesting.

    • Forecasting
    • Gradient boosting
    • Backtesting
  • british-invoice-digitization

    Computer vision for invoice field extraction with YOLOv5 and a FastAPI service.

    • Computer vision
    • YOLOv5
    • FastAPI

Standards across these repositories

  • Each project is a tested Python package with continuous integration, and a docs/WHY.md that explains its design decisions.
  • Original shell scripts sit in legacy/ beside the rewrites, to show how ad-hoc tooling became maintainable software.
  • No real participant data, credentials or proprietary information. Synthetic data only.
  • I build with AI coding assistants and say so: commits written with Claude carry a co-author line. I review, test and own everything that is merged.