Trivedi Centre for Political Data (TCPD)
At TCPD, I worked at the intersection of software engineering and political science, building production data systems that transformed fragmented government records into research-ready datasets. My work ranged from ETL pipelines and researcher-facing web applications to production infrastructure and large-scale social media datasets, enabling researchers to study Indian politics at scale.
Selected technical highlights
- Production ETL pipeline engineering
- Historical PDF & web data extraction
- Full-stack Django application development
- Research infrastructure (AWS, Docker, NGINX)
- Large-scale data quality & annotation workflows
Tech Stack
Python, Pandas, Django, JavaScript, AWS LightSail, Docker, NGINX, Tabula, Regex
About TCPD
The Trivedi Centre for Political Data (TCPD) at Ashoka University was a research centre dedicated to building high-quality datasets on Indian political life. Raw election and legislative data published by the Indian government is fragmented, inconsistent, and often locked in PDFs or poorly structured files, making it impossible to use for analysis. TCPD was founded to bridge this gap.
As the centre’s primary research engineer, I maintained production systems, rebuilt core data pipelines, developed researcher-facing software, and collaborated closely with political scientists to translate research workflows into reliable engineering systems
TCPD was dissolved in 2023.
The data is now hosted under the Centre for Data Science and Analytics, Ashoka University.
My Role
I joined TCPD as a Research Engineering Intern in December 2021 and transitioned into a full-time Research Engineer the following year. Over time, I took ownership of much of the centre’s engineering infrastructure, maintaining production systems, building new research tools, and supporting multiple interdisciplinary projects.
Key Contributions
My work combined data engineering, software development, and research collaboration across several long-term projects. While each project addressed different research questions, they shared a common goal: building reliable systems that transformed messy public data into reusable research infrastructure.
1. Lok Dhaba
Lok Dhaba was TCPD’s flagship dataset: a repository of Indian election results, widely used by researchers, journalists, and policy analysts. In addition to outcomes, it included candidate-level details such as education, profession, and unique IDs to track political careers over time.
My contributions included:
- Rewriting an undocumented ~1,500-line R ETL pipeline into a modular Python workflow, improving reliability, maintainability, validation, and fault tolerance.
- Extending the Indian Elections Dataset from 1962 back to India's first general election (1951) by building extraction pipelines for 50+ historical Election Commission PDFs.
- Handling historical complexities including multi-member constituencies, changing state boundaries, and reserved seats through custom parsing and validation workflows.
- Maintaining Lok Dhaba's production pipelines and public data infrastructure while supporting downstream datasets such as the Political Career Tracker.
2. Collaborative Research Platform
One of TCPD’s research projects involved manually annotating candidate records using evidence gathered from diverse public sources. The annotations often required subjective judgement and evolved as new information became available, making transparency, reviewability, and reproducibility central to the research process.
To support this workflow, I designed and developed an internal web application that replaced spreadsheet-based collaboration with a centralized, version-controlled annotation system.
Key features included:
- A Django-based web application deployed on AWS Lightsail for managing collaborative research annotation workflows.
- Authentication and role-based access, allowing multiple researchers to work simultaneously while maintaining data integrity.
- Version-controlled annotations with complete audit history, enabling researchers to review previous decisions, track revisions over time, and reproduce coding choices.
- Approval workflows and researcher-friendly interfaces for searching records, recording supporting evidence, and monitoring project progress.
The platform transformed a fragmented manual workflow into a structured research system, improving collaboration, consistency, and reproducibility while providing a transparent record of how annotation decisions evolved over time.
2. Social Media Project (Meta, Twitter)
This project tracked political advertising on Meta across multiple state elections, providing one of the first systematic datasets on digital election campaigning in India.
My contributions included:
- Building data pipelines around Meta’s Marketing API to collect, clean, and standardize political advertising data across multiple state elections.
- Designing annotation methodologies for ad content and implementing quality-control workflows to ensure consistent coding across researchers.
- Linking advertisers, candidates, constituencies, and election outcomes to enable downstream analyses of campaign strategies and spending patterns.
- Producing exploratory analyses and visualisations that were presented at the Social Media and Society Conference (University of Michigan, 2022) and informed two published Hindustan Times articles on digital election campaigning.
The project demonstrated how computational methods could be used to systematically study political advertising at scale, offering new insights into campaign strategies across Indian elections.
3. Digital Society Project (DSP)
The Digital Society Project (DSP) studies the online presence of political actors worldwide. TCPD collaborated with DSP to build a comprehensive dataset of Indian politicians’ Twitter presence during election cycles, published here.
My responsibilities included:
- Leading TCPD’s collaboration with DSP, serving as the primary point of contact with external researchers and coordinating project delivery.
- Managing teams of annotators responsible for identifying, validating, and documenting politicians’ official Twitter accounts across multiple elections.
- Building data pipelines that integrated Lok Dhaba, cabinet, and Twitter datasets into DSP’s publication schema while implementing validation and quality checks.
- Collecting and processing over 100,000 tweets, documenting data limitations, and delivering cleaned datasets aligned with election timelines.
This collaboration expanded TCPD’s research beyond Meta advertising, creating one of the most comprehensive publicly available datasets on Indian politicians’ social media presence.
Research Infrastructure
Beyond individual datasets, I was responsible for maintaining the engineering infrastructure supporting TCPD’s public data services.
This included:
- Deploying and maintaining applications on AWS Lightsail using Docker and NGINX.
- Improving public dataset delivery by replacing database-backed downloads with pre-generated archives served directly through NGINX.
- Supporting researchers with internal tools, deployments, and production data workflows.
Publications
News Articles
- Hindustan Times / Dec’ 12, 2022: “Number Theory: How Parties used Social Media in Gujarat Elections”
- Hindustan Times / Mar’ 22, 2022: Punjab Election: How Candidates and Parties used Social Media
Dataset Contributions
- “TCPD Indian Elections dataset (TCPD-IED), 1951-1962”. Trivedi Centre for Political Data, Ashoka University.
- “TCPD Rajya Sabha dataset (TCPD-RSD), 1952 – 2022”. Trivedi Centre for Political Data, Ashoka University.
- “TCPD Judiciary dataset (TCPD-IJD), 1950–2021”. Trivedi Centre for Political Data, Ashoka University.
- “TCPD and DSP Indian Politicians’ Social Media (TCPD-DSP IPSM) Dataset.” Digital Society Project.