Data Science Associate

About the Role

ARTPARK, the AI & Robotics Technology Park at the Indian Institute of Science (IISc), is a hub for translational research focused on creating significant social impact. We pioneer solutions by bringing together data scientists, researchers, policymakers, and technology startups.

ARTPARK’s Data Science team works with national and state institutions on livestock health and production. This role offers the opportunity to work with large, real-world datasets, including national livestock censuses, production and yield statistics, household sample surveys, disease and environmental surveillance records, and geospatial data, to produce analysis that supports animal health and livestock development programmes. You will collaborate with government stakeholders, veterinary and animal husbandry departments, engineers, and disease modelling experts, as well as research partners including the National Centre for Biological Sciences (NCBS) and Penn State University, in a national and international research setting.

This position is open to both fresh graduates and candidates with relevant professional experience. Fresh graduates are encouraged to apply.

Location: Bengaluru

Roles & Responsibilities

Data Preparation: Own the end-to-end lifecycle of livestock and animal health data: acquire, clean, harmonise, and link data from government sources, surveys, and surveillance systems into analysis-ready datasets.

Analysis: Analyse livestock population, production, yield, and animal health datasets across species, geography, and time to identify patterns, trends, and differences between regions.

Methodology and Documentation: Work with the team’s library of standard operating procedures, the written methods that define how each recurring type of question is answered. Apply them, improve them, and write new ones, stating clearly what data each method needs and what its limitations are.

Modelling and Machine Learning: Build and validate statistical and machine learning models on real-world data. This includes feature engineering, selecting a model suited to the question, and testing it honestly on held-out data.

Data Pipelines: Contribute to the pipelines that ingest, clean, and publish these datasets. Turn one-off analysis scripts into repeatable steps with automated quality checks, working alongside the engineering team.

Quality Control: Check results before they are used. Verify totals, units, and agreement between independent sources, and investigate discrepancies rather than smoothing over them.

Communication and Reporting: Translate analytical findings into clear reports, charts, maps, and presentations for programme staff, veterinarians, government officials, and research collaborators.

Collaboration: Work alongside engineers, epidemiologists, modellers, and domain experts to refine methods, review each other’s work, and share knowledge.


Experience & Qualifications

Education: A Bachelor’s or Master’s in statistics, data science, computer science, mathematics, economics, agricultural or animal sciences (quantitative), public health (quantitative), or a related quantitative discipline.

Programming Skills: Working proficiency in Python (pandas, numpy) or R (tidyverse) for data manipulation and analysis, gained through coursework, projects, or internships. Familiarity with version control (git).

Statistical Foundations: Solid grounding in applied statistics, including descriptive statistics, rates and ratios, growth rates, regression, and the basics of time-series analysis. 

Machine Learning: Practical experience building models on tabular data, covering feature engineering, model selection, and validation, with awareness of common pitfalls such as data leakage, overfitting, and class imbalance.

Data Handling: Comfort working with messy, incomplete, real-world data: missing values, inconsistent units, revised or duplicated records, and combining sources that do not line up cleanly.

Data Engineering: SQL, and the ability to structure code as a repeatable pipeline rather than a script run by hand. Working knowledge of cloud platforms (AWS, GCP, or Azure) is a plus.

Geospatial Data: Some exposure to working with maps or location-based data, and comfort with data reported by state and district.

Communication: Ability to explain methods and findings in plain language to non-technical audiences, particularly government and veterinary stakeholders.

Rigour and Honesty: Care in checking your own numbers, and a willingness to say when the data cannot answer a question rather than presenting a result that overstates what is known.

Growth Mindset: A willingness to learn a new domain, engage with new challenges, and quickly become proficient with new tools and methods.

Desired Experience: Prior exposure to livestock, agriculture, veterinary, animal health, or environmental surveillance data is desirable. Experience with census data, large government surveys, or GIS and spatial analysis is a strong plus, as is experience working with or within government institutions.


ARTPARK @ IISc :  Innovation factory for next-gen robotics & AI
ARTPARK is India's leading deep-tech venture builder and incubator focused on robotics, connected autonomous systems, and AI. Leveraging our unique facilities and ecosystems, we strive to provide meaningful support to very early-stage startups building deep-tech products based in research. We are a nonprofit organization created by Indian Institute of Science (IISc, Bengaluru) with support from the Department of Science & Technology (Government of India) and the Government of Karnataka.

Previous
Previous

AI Engineer

Next
Next

Data Science Intern