Senior AI Alignment Specialist
Open to offersNew to PlatformSandeep D. is a seasoned AI Alignment & Multilingual Model Training Specialist based in Varanasi, India, with extensive hands-on experience in Reinforcement Learning from Human Feedback (RLHF) and Supervised Fine-Tuning (SFT). With over 16 years in linguistic and editorial realms, Sandeep artfully employs cultural-linguistic reasoning to enhance semantic alignment and instruction adherence. His expertise in hallucination & bias detection and red teaming ensures robust safety evaluations and cultural context precision across Hindi-English datasets. At Outlier.ai, he crafted bilingual prompts for semantic and contextual accuracy in multilingual AI systems, excelling in structured preference ranking and factual evaluations. His tenure includes pivotal roles at CrowdGen and Alignerr, focusing on AI content quality and multimodal evaluation. Sandeep's academic pursuits include a Bachelor in Tourism Studies and diplomas in Computer Applications and Urdu Language, complementing his professional mastery in multilingual model alignment.
Sign in as an employer to save this profile or invite Sandeep D. to a job.
Sign in as an employer