Data footprint
Large-scale unstructured data preparation: cleaning, transformation, and pipeline work that turns raw records into a usable training and evaluation surface.
Deep dive
The delivery layer behind model work: Databricks workflows, employee-facing RAG, data preparation, annotation and synthetic-data throughput, experiment iteration, and maintainable code.
Strong model work depends on the layer behind it: data, tooling, governed release controls, team coordination, and repeatable iteration.
Data footprint
Large-scale unstructured data preparation: cleaning, transformation, and pipeline work that turns raw records into a usable training and evaluation surface.
Platform layer
Employee-facing retrieval assistant on governed platform services, with answer quality raised through HyDE, multi-query retrieval, and reranking, then checked against human-verified evaluation sets.
How the work runs
Release gates, audit logging, token/cost tracking, annotation automation, synthetic coverage, refactoring, and team coordination move together.
Platform map
The path moves through intake, preparation, annotation or augmentation, model iteration, and operational feedback.
01
Work starts with large structured and unstructured sources, not clean benchmark inputs.
02
Cleaning, transformation, Databricks-backed handling, and release controls create a usable training and evaluation surface.
03
Automation and synthetic data expansion help reduce manual-label bottlenecks in class-heavy NLP work.
04
Transformers, LLMs, RAG, and supporting experiments move faster when the data layer advances with them.
05
Refactoring, mentoring, token/cost tracking, and small-team coordination keep the work maintainable.
Delivery layers
Databricks workflows
Recent delivery pairs LLM and NLP work with Azure Databricks so large-scale data handling and experimentation can move inside the same operating surface.
Data preparation
Preparing unstructured records at scale, then replacing per-experiment setup with config-driven preprocessing and embedding pipelines, shows depth beyond clean benchmark data.
Annotation automation
Automated collection, annotation support, and synthetic-data expansion reduced labeling pressure in a many-class NLP problem and kept experimentation moving.
Mentoring
Refactoring Python code, mentoring junior data scientists, and managing 4 direct-report data science interns at Blue Guardian helped turn model work into a maintained delivery path.
Results
Data preparation
Unstructured data at scale
Prepared for generative chemistry model work at Servier Canada.
Employee RAG
Measured retrieval quality
Employee-facing retrieval assistant on governed platform services, scored with human-verified evaluation sets and LLM-as-judge, then piloted with support staff.
Model preparation
Config-driven preparation
Preprocessing and embedding pipelines moved into configuration, so a new model run starts from a prepared surface instead of a fresh script.
Annotation cost
Manual labeling displaced
LLM-based data generation, validation, and automated annotation absorbed labeling work that would otherwise have been done by hand.
Coverage expansion
LLM-generated synthetic text
Used to expand class coverage and accelerate iteration in multi-class mental-health text classification.
Forms workflow POC
Evaluated against historical cases
Report-drafting agents with human review, aimed at the rework loop that starts when a form comes back for modification, and evaluated end to end on historical cases rather than a demo.
Agentic maintenance
Failing runs triaged automatically
An agent harness diagnoses a failed pipeline run, validates the fix on sandboxed data samples, and stages a deploy-ready change for human approval. Later extended to automated code review.
Claim-duration model
Simplest model that met the bar
A gradient-boosted claim-duration classifier chosen over transformer alternatives after benchmarking accuracy, latency, and cost, then kept under drift monitoring with retraining alerts.
What makes it real
01
The data path is maintained, not treated as a one-time preprocessing script.
02
Automation is used to reduce bottlenecks, not to hide weak workflow ownership.
03
Leadership shows up as iteration speed, cleaner code, and team throughput.
The practical evidence is in the data preparation, platform decisions, automation strategy, and delivery loop that keep experimentation grounded.