Data Curation Associate
Type: Full time On-site
Employment Type: Contractual
No. of Position: 02
About the Team & Project
You will be a vital part of the Language Data & AI team, focused on building a robust ecosystem for AI in Indian languages. Our goal is to create high-quality, open-source speech and text datasets spanning various districts to accelerate the state-of-the-art in Natural Language Processing (NLP) for societal impact. As part of this ambitious, nationwide program, you will join the operations team responsible for driving data collection and curation across major language initiatives. You will work closely with the core team to spearhead AI advancements in Indic languages through speech data collection, curation, and advanced language modeling.
Note: Each selected candidate will be required to manage operations and take ownership of two or more states.
Role & Key Responsibilities
You will be responsible for overseeing and managing data curation from one or more of the following regions:
● Uttar Pradesh: Hindi, Bhojpuri, Awadhi, Urdu, Khariboli, Bhatri, Dorli, Kurumali, Duruwa, Bantar etc- (primarily from Budaun, Etah, Varanasi, Jyotiba Phule Nagar, Gorakhpur, Hamirpur, Ghazipur, Muzaffarnagar, Deoria, etc)
● Bihar: Hindi, Maithili, Magahi, Bhojpuri Angika, Bajjika etc- (primarily from Patna, Gaya, Muzaffarpur, Darbhanga, Saharsa, Supaul, Bhagalpur, Begusarai, Purnia, etc)
● Maharashtra: Marathi- (primarily from Washim, Gondia, Mumbai suburban, Pune, Nagpur, Chandrapur, Dhule, Solapur, Aurangabad, Sindhudurg etc)
● Goa: Konkani, Marathi- (primarily from North Goa, South Goa etc)
Your day-to-day responsibilities will include:
● Data Curation Management: Understand the specific requirements for data curation and how they impact the AI models being built. Manage daily curation operations foraudio and transcription data.
● Recruitment & Outreach: Search, identify, and recruit the right curation experts through local contacts, NGOs, and educational institutes in your assigned districts or languages. Design outreach flyers and contact potential applicants (via phone and WhatsApp) to drive interest in data curation and quality checking.
● Onboarding & Training: Host awareness calls with potential candidates to help them understand the tasks. Train the recruited experts and provide them with all relevant guideline documentation.
● Supervision & Workflow: Assign daily workloads based on the experts’ availability. Continuously supervise, coordinate, and review their daily performance to ensure quality and timely delivery of work.
Candidate Background & Requirements
● Language Skills: Must be a native speaker of the local languages/dialects prominent in the target states mentioned above.
● Communication: Excellent verbal and written communication skills in both English and the relevant local language.
● Team Coordination: Demonstrated ability to handle, motivate, and manage multiple people working remotely to get tasks completed efficiently.
● Technical Proficiency: Strong working knowledge of Microsoft Office (Excel, Word, PowerPoint) and Google Workspace (Docs, Sheets, Slides).
Preferred (Good-to-Have) Skills
● Prior experience in data curation, data sourcing, or working with annotation companies.
● Hands-on experience with speech data annotation and labeling.
ARTPARK @ IISc : Innovation factory for next-gen robotics & AI
ARTPARK is India's leading deep-tech venture builder and incubator focused on robotics, connected autonomous systems, and AI. Leveraging our unique facilities and ecosystems, we strive to provide meaningful support to very early-stage startups building deep-tech products based in research. We are a nonprofit organization created by Indian Institute of Science (IISc, Bengaluru) with support from the Department of Science & Technology (Government of India) and the Government of Karnataka.