Title: ML Engineer, Training Pipelines (ML Infrastructure)
Company Name: Labaid AI Ltd.
Vacancy: 2
Age: At least 22 years
Job Location: Dhaka
Salary: Negotiable
Experience:
Published: 2026-07-19
Application Deadline: 2026-07-31
Education:
Requirements:
Skills Required:
Additional Requirements:

ML Engineer, Training Pipelines (ML Infrastructure)
Company:
Labaid AI (a Labaid Group company)
Position:
Machine Learning Engineer — Training Pipelines & ML Infrastructure
Vacancy:
01
Job Location:
Dhaka
Employment Status:
Full-time
Salary:
Negotiable
About Labaid AI
Labaid AI is building trusted vertical AI for healthcare, identity, and edge intelligence in Bangladesh, backed by Labaid Group's clinical network (~3 million annual patient encounters). Our product family — LUNA (agentic healthcare assistant platform with a medical LLM + RAG core), MedPAC (AI-assisted enterprise imaging), MyHealth (oncology copilot), Digital RM (AI relationship manager), and FLVE / Edge Vision (liveness and edge AI appliances) — depends on models that are continuously trained, fine-tuned, evaluated, and shipped. This role owns the machinery that makes that possible.
Job Context
We are hiring an ML Engineer to design and own our end-to-end model training pipeline: from raw clinical/imaging/text data to fine-tuned, evaluated, optimized models running in cloud, on-prem hospital, and edge environments. Most local "ML Engineer" roles are model-user roles; this one is a model-factory role. You will build the infrastructure that lets our data scientists and researchers train medical LLMs, Bangla NLP models, and medical vision models reliably and reproducibly.
Job Responsibilities
· Design, build, and operate training and fine-tuning pipelines for:
· LLMs — supervised fine-tuning, LoRA/QLoRA/PEFT, instruction tuning, and preference tuning of medical and Bangla/Banglish language models; RAG index build pipelines (embeddings, chunking, refresh).
· Vision models — classification, detection, and segmentation for medical imaging (MedPAC) and edge vision (StockSight/SecureSight-class products); face liveness/anti-spoofing models (FLVE).
· Speech/NLP — Bangla ASR and text pipelines as we expand voice interfaces.
· Build the data side of training: ingestion from hospital systems, de-identification/PHI-scrubbing stages, dataset versioning (DVC/lakeFS or equivalent), annotation pipeline integration, and dataset quality gates.
· Stand up and maintain experiment tracking and model registry (MLflow or Weights & Biases): every model reproducible from config + data version + code commit.
· Manage GPU training infrastructure: job scheduling, multi-GPU/distributed training (DeepSpeed/FSDP/Accelerate), utilization monitoring, cost control across on-prem GPUs and cloud burst capacity.
· Own model optimization and packaging for deployment: quantization, ONNX export, TensorRT/OpenVINO optimization, and handoff to serving (Triton/vLLM) across cloud, on-prem hospital servers, and edge devices — the same model must ship to a data centre and a low-power appliance.
· Build evaluation into the pipeline: automated benchmark suites (clinical accuracy, Bangla quality, safety/refusal behaviour, imaging metrics), regression detection before any model is promoted, and post-deployment drift monitoring feeding back into retraining.
· Establish CI/CD for models: automated retraining triggers, canary promotion, rollback, and audit trails suitable for healthcare and regulated (eKYC/BFSI) buyers.
· Write clean, tested pipeline code and documentation; champion reproducible ML practice across the team.
Educational Requirements
· B.Sc. in Computer Science & Engineering or a related field from a reputed university; M.Sc. is a plus.
· Strong open-source contributions or a demonstrable track record of shipped ML systems are weighted alongside degrees.
Experience Requirements
· 1–3 years in software/ML engineering, with hands-on experience training or deploying deep learning models.
· Candidates from fintech, telco big-data teams, AI startups, or research labs are all welcome; healthcare experience is a plus, not a requirement.
Additional Requirements (Skills)
Must have:
· Excellent Python engineering (not just notebooks): packaging, testing, typing, code review discipline.
· PyTorch and the Hugging Face ecosystem (Transformers, Datasets, PEFT/TRL).
· Docker and Linux fluency; comfort operating GPU servers (CUDA, drivers, NVML monitoring).
· Pipeline orchestration experience (Airflow, Prefect, Dagster, Kubeflow, or equivalent — tool matters less than the discipline).
· Experiment tracking and reproducibility practice (MLflow, W&B, DVC, or equivalent).
· SQL and comfort with data at scale.
Nice to have:
· Distributed training (DeepSpeed, FSDP, Ray) and LLM fine-tuning at 7B+ scale.
· Inference optimization: ONNX Runtime, TensorRT, OpenVINO, Triton, vLLM, quantization (GPTQ/AWQ/INT8).
· Kubernetes; AWS or GCP (we run hybrid: cloud + on-prem hospital + edge).
· Vector database operations (Qdrant, pgvector, FAISS).
· Bangla NLP/ASR datasets and tooling.
· Medical imaging formats (DICOM) or edge hardware (Jetson-class devices).
Other:
· Both males and females are encouraged to apply.
· No age limit — we hire for capability.
What You Get
· Competitive salary, reviewed yearly.
· 2 festival bonuses per year (as per company policy).
· Medical coverage through the Labaid healthcare network.
· Mobile allowance.
· Weekly holiday: Friday.
· Real GPU infrastructure to run and grow — you will own it, not queue for it.
· One of the very few roles in Bangladesh training models on proprietary clinical, imaging, and Bangla medical data.
· Direct impact: models you ship reach hospitals, cancer centres, banks, and factories across the country.
Interview Format — Come Prepared to Showcase Your Work
· Project showcase (mandatory): Present your own work covering both pipelines and data science — the engineering side (data/training pipelines, orchestration, deployment, monitoring, what broke and how you fixed it) and the modeling side (what was trained, how it was evaluated, what the metrics actually meant). Bring your laptop with code, configs, or a live demo. Slides alone are not enough — we will ask to see the work.
· Technical deep-dive Q&A: Expect detailed questions on your projects and on fundamentals — Python engineering, ML training mechanics, data handling, reproducibility, deployment, and enough statistics/ML theory to reason about what your pipelines produce. Be ready to defend every decision: why this architecture, why this tool, what failed, what you'd do differently.
· We are assessing understanding, not memorization. Candidates who can clearly explain the why behind their work will stand out over those with impressive-sounding but shallow project lists.
How to Apply
Email your CV (PDF) to hr@labaidcancer.com with subject line: ML Training Pipeline Engineer – [Your Name].
Include your GitHub and a short description of the most complex training or data pipeline you have built and operated — what broke, and how you made it reliable.
Labaid AI is an equal-opportunity employer. All model training on clinical data follows strict consent, de-identification, and audit policies.