I build multimodal ML systems that run at internet scale.
Founding member of IAS's $150M trust & safety classification suite. Detecting low-prevalence unsafe content across 3 Trillion impressions annually. Patented.
Real systems running in production at scale.
Founding member of the Multimodal ML team. Built the video harm detection system covering 12 GARM unsafe-content and 37 contextual categories across TikTok, YouTube, Meta, and X. Processes 50 years of video per day and 3 trillion impressions per year, powering ~$150M in annual revenue (~30% of company). Patented (US 12,217,493).
Ideated, prototyped and shipped ALP, an automated labeling platform: text, image and video embeddings, vector index, ANN search, and VLM/LLM review of positive matches. Took it from proof of concept to a company-wide, cross-team platform. Cut labeling cost and turnaround by 99.5%.
Led PEFT/LoRA adoption for deepfake detection models, slashing compute by ~80% and boosting experiment throughput 3x. No performance degradation.
Took open-source decision models from idea to production in one week. Benchmarked variants, scaled scoring on AWS Auto Scaling Groups, and pitched the CTO and leadership on placing it in front of our classifiers, cutting an estimated ~70% of costs on production data.
Directed end-to-end. Prototyped a teacher model using open-source LLMs with RAG to label training data, then trained and productionized DistilBERT models extended for long token lengths to flag misleading narratives across 22M+ videos per day.
Built a pseudo-labeling pipeline using SigLIP and Vicuna that cut model development cost by 90%+ across 12 unsafe categories. Separately led synthetic Twitter data generation with GPT-Neo that raised harmful-content classifier precision by 49%.
Built custom translation models with OpenNMT for 42 language pairs. Translates 75 trillion characters annually at 67% lower cost, enabling multilingual harm detection in production pipelines.
Built a causal inference pipeline ingesting 8 billion ad impressions per day to measure real cause-and-effect on conversion rates for ad campaigns using Bayesian methods.
Built predictive models and churn strategies that identified customer behavior patterns and converted 18% of one-time buyers into repeat customers.
We classify extremely low-prevalence, nuanced categories of unsafe content in social media video. Previously, building a new classifier meant weeks of work: sampling the right balance of rare data from a large amount of production data which is actually just safe content, designing experiments to extract signal, then paying for expensive, time-consuming human annotations on subjective user-generated video. Each iteration was slow, costly, and hard to scale.
An internal Python package that closes the entire loop. It finds relevant content in production via vector search, accepts custom prompts (versioned in GitHub), and processes any modality. For video, it extracts keyframes, deduplicates frames, then sends everything to a cost-optimized LLM + VLM system for multimodal fusion. The output includes multilabel classifications, topic/subtopic stratification for comprehensive dataset sampling, and artifacts ready for model training or human review.
Human reviewers validate AI labels through a custom A/B testing UI. Their responses feed back into the package, where DSPy and GEPA optimize the original prompts automatically. This is essentially an RLAIF pipeline. The LLM-as-judge acts as a reward model, with engineered reward functions calibrated against human baselines to prevent reward hacking. The loop runs continuously: label, validate, optimize, retrain.
Not a laundry list. Every skill tied to something I shipped.
Open to conversations about ML, data science leadership, and impactful opportunities.