Daniyal Khan

AI Research Engineer

Daniyal Khan

I turn unstructured speech and text into production systems — speech processing, NLP and generative AI, taken from research prototype to deployed product.

Daniyal Khan, AI Research Engineer

Selected Work

Things I've built

All work →

Voice Fingerprinting Pipeline

Isolating individual customer conversations inside eight-hour, multi-speaker audio recorded on wearables across retail stores.

~60%
Precision
~49%
Recall
100k+ hrs
Audio Processed
  • Python
  • PyTorch
  • Silero VAD
  • Pyannote
  • ReDimNet
  • Qdrant
  • Resemble AI
Read case study →

RAG-Based Event Detection Pipeline

Finding where one customer interaction ends and the next begins, by classifying what happened rather than who was speaking.

+59%
Customer Coverage
66%
Personal Speech Cut
2,000+
Stores Live
  • Python
  • Gemini
  • Qwen Embeddings
  • RAPTOR
  • HyDE
  • KNN
  • Qdrant
Read case study →

Diabetic Retinopathy Diagnosis with Interpretable CNNs

Screening retinal images for diabetic retinopathy severity, with visual explanations a clinician can actually audit.

0.82
Quadratic Kappa
5
Severity Classes
Grad-CAM + LIME
Explainability
  • Python
  • TensorFlow
  • InceptionV3
  • Grad-CAM
  • LIME
  • OpenCV
Read case study →

Multimodal Sentiment Analysis with Tensor Fusion

Reading sentiment from speech and text together, because tone routinely contradicts the words.

Audio + Text
Modalities
Tensor Fusion
Fusion
Transformer
Architecture
  • Python
  • PyTorch
  • Transformers
  • Librosa
  • melSpectrogram
Read case study →

Approach

How I build ML systems

Problem

Write down the cost of being wrong before writing any code.

What decision does this model inform, and what does each kind of error actually cost? A false positive and a false negative are almost never equally expensive, and that asymmetry should drive the metric, the threshold and sometimes the entire approach.

  • Problem framing
  • Success criteria

Writing

Recent notes

All notes →

Accuracy Is Usually the Wrong Target

On ordinal labels, imbalanced classes, and why the metric you optimise is a modelling decision rather than a reporting one.

  • #evaluation
  • #metrics

Experience

Where I've worked

AI Research Engineer

Aug 2024 — Present

YoyoAI · Bengaluru

Extracting customer–salesperson interactions from 8-hour multi-speaker audio captured on wearable devices across retail stores.

  • Built a voice fingerprinting pipeline — denoising, VAD, speaker diarization, embedding generation and custom clustering — with Qdrant-backed retrieval for cross-session speaker tracking.
  • Designed a RAG-style event detection pipeline using HyDE multi-query prompting and RAPTOR-inspired clustering, lifting customer coverage by 59% and cutting personal speech by 66%.
  • Combined both pipelines with LLM-based filtering to reach 89% precision and 73% recall on production audio.
  • Scaled the system across 100,000+ hours of audio and from 2 to 8 clients covering 2,000+ retail stores in India, taking it from research prototype to deployed product.

Toolkit

What I work with

Languages & Frameworks

  • Python
  • SQL
  • PyTorch
  • TensorFlow

Machine Learning & AI

  • Deep Learning
  • NLP
  • Computer Vision
  • Generative AI
  • Classical ML
  • Statistical Modeling

Speech & Audio

  • Speaker Diarization
  • VAD
  • Speaker Embeddings
  • Denoising

Experimentation & Tracking

  • DVC
  • ClearML
  • Weights & Biases

MLOps

  • Docker
  • Kubernetes
  • AWS
  • CI/CD
  • FastAPI
  • Flask
  • BentoML

Data & Retrieval

  • Qdrant
  • RAG
  • Vector Search
  • REST APIs
  • Git

Contact

Get in touch

Open to conversations about machine learning engineering roles, interesting problems, or anything on this site you'd like to dig into.