Skip to content
Open to opportunities

Sahil Pandita

Lead Data Engineer — Big Data Platforms, ML Data Pipelines & AI-Powered Data Tools

12 years building production data systems at telecom scale. Currently at Airtel X Labs, where my work supports India's first AI-powered spam detection solution, protecting 300M+ users.

SOURCESPROCESSSERVEOracleKafka / SolaceCDC · GoldenGateSparkAb InitioHiveAerospikeML modelAI agent

01 / about

About

I'm a data engineer who has spent 12 years turning messy, high-volume telecom and enterprise data into reliable, governed systems that are ready for analytics and ML. I've led migrations from Oracle to Hadoop data lakes and built CDC pipelines from GoldenGate and Solace queues into Hive. Today I lead the DARTS BI reporting team at Airtel X Labs and work on the data platform behind Airtel's AI-powered spam detection.

Lately I've been combining that foundation with AI. I completed a PGP in AI & ML from Great Lakes and UT Austin. I also built a PySpark framework that monitors data drift across ML pipelines, and an LLM-powered agent that lets my team query spam-detection data in plain English. I care about reliability, about automation that removes manual toil, and about data you can trust.

years in data engineering
12+
telecom users protected by the spam platform I work on
300M+
records × features monitored for drift per run
600M+ × ~70
automated tests passing on my LLM data agent
202/202
companies: Airtel, Capgemini, Exusia, Accenture
4

02 / experience

Experience

  1. Lead Engineer — Airtel X Labs

    Dec 2021 – Present · 4 yrs 11 mos

    Airtel Digital India Pvt. Ltd. · India

    • Lead engineer for DARTS BI Reporting, managing and mentoring a team that delivers business-critical analytics.
    • Played a key role in India's first AI-powered spam detection solution, protecting 300M+ telecom users from spam and fraud; contributed to system architecture and integration with telecom networks.
    • Designed and built a scalable, reusable PySpark framework that detects and monitors data drift across multiple ML pipelines and raises early alerts when data changes.
    • PySpark
    • Spark on YARN
    • Hive
    • HDFS
    • Alluxio
    • Ab Initio
    • Aerospike
    • Airflow
    • GoldenGate CDC
    • Solace
    • GCP
    • AWS S3
    • Python
    • LLMs
  2. Software Development Consultant — Capgemini India

    Aug 2018 – Dec 2021 · 3 yrs 5 mos

    Pune

    • Led a data lake migration, extracting data from Oracle source systems into Hadoop-based data lakes.
    • Built utilities to automate DML creation and test-data generation, cutting manual effort and errors.
    • Designed batch processes that analysed every migrated table and produced validation reports.
    • Ab Initio
    • Hadoop
    • Hive
    • Oracle
    • CDC
    • Unix shell
  3. Senior Analyst — Exusia

    Jul 2017 – Aug 2018 · 1 yr 2 mos

    Pune

    • Executed the PAMS data migration from Oracle (via stored procedures) into Netezza (COGS).
    • Built parameter sets (psets) with Express>It inside an in-house Ab Initio framework.
    • Ab Initio
    • Express>It
    • Oracle
    • Netezza
    • PL/SQL
  4. Accenture

    Aug 2014 – Jun 2017 · 2 yrs 11 mos

    Gurugram

    • Application Development AnalystDec 2015 – Jun 2017 · 1 yr 7 mos
    • Associate Software EngineerAug 2014 – Dec 2015 · 1 yr 5 mos
    • Worked on Global Data Quality programmes to analyse, measure and report on critical data elements; authored data quality rules.
    • Migrated large-scale DQ and EVM projects from AIX to Linux with zero data loss.
    • Built an Ab Initio regression utility comparing AIX and Linux outputs and table loads, generating Excel migration reports.
    • Ab Initio
    • Data Quality
    • AIX
    • Linux
    • Unix shell

03 / projects

Projects

Selected work across AI tooling, ML data platforms, pipelines and automation. Each has a short case study.

Showing 12 projects

  • AI / LLMFeaturedAirtel X Labs

    AI-Powered Natural-Language Data Agent

    A conversational agent that turns plain-English questions into Hive SQL, runs them on the cluster and replies with results, charts or files.

    automated tests passing
    202/202
    pass rate when the harness started
    35/44
    • Python
    • PySpark
    • Hive
    • HDFS
    • YARN
    • LLMs
    • Gemma
    • Prompt Engineering
    • Telegram API
    • Confluence
  • ML Data PlatformFeaturedAirtel X Labs

    Data Drift Monitoring Framework for ML Pipelines

    A reusable PySpark framework that compares feature distributions between time windows and raises early alerts, across multiple spam ML pipelines.

    records per run
    600M+
    features tested for drift
    ~70
    • PySpark
    • Spark on YARN
    • Statistics
    • MLOps
    • Parquet
    • HDFS
  • ML Data PlatformFeaturedAirtel X Labs

    Spam Detection Data Platform

    The data platform behind India's first AI-powered spam detection solution, protecting 300M+ Airtel users from spam and fraud calls and SMS.

    Airtel users protected
    300M+
    AI-powered spam detection in India
    1st
    • PySpark
    • Hive
    • Aerospike
    • Airflow
    • YARN
    • Ab Initio
    • ML features
  • Data Pipelines & MigrationAirtel X Labs

    Siebel CDC Reporting Pipeline

    A continuous change-data-capture pipeline from Oracle GoldenGate and Solace queues into Hive, feeding daily postpaid and Telemedia reports.

    • CDC
    • GoldenGate
    • Solace
    • Hive
    • Ab Initio
  • Data Pipelines & MigrationAirtel X Labs

    GCP Reporting & SMS Insights

    Ab Initio pipelines that reintegrate reports into GCP, and automated SMS Insights reporting on the bulk SMS aggregator market, delivered to AWS S3.

    • Ab Initio
    • GCP
    • AWS S3
    • SQL
  • Data Pipelines & MigrationCapgemini

    Oracle Data Lake Migration

    Led the migration of Oracle source systems into a Hadoop data lake, with generated DMLs, test data and table-by-table validation.

    • Ab Initio
    • Oracle
    • Hadoop
    • Hive
    • Unix shell
  • Data Pipelines & Migration

    Oracle Hierarchical Logic → PySpark

    Re-implemented an Oracle PL/SQL account-hierarchy function (CONNECT BY) as a distributed, iterative PySpark job with checkpointing.

    • PySpark
    • Oracle PL/SQL
    • Spark SQL
  • Automation & Reliability

    Ab Initio Attribute Onboarding Automation

    Automates onboarding new business attributes into multiple Ab Initio DMLs, an XFR and a JSON config, with developer approval and auto-tagging.

    • Ab Initio
    • EME/air commands
    • Shell
    • Python
  • Automation & Reliability

    Self-Healing Metadata Watchdog for Hive/Alluxio

    Detects stale snapshot tables, auto-remediates Hive/HDFS metadata drift, and alerts only when upstream data is genuinely missing.

    • Bash
    • Hive
    • Alluxio
    • HDFS
    • Monitoring
  • Automation & ReliabilityAirtel X Labs

    Enterprise PII Encryption & Reporting Standardisation

    PII data encryption across the DARTS ecosystem, and the AIO Reporting framework to standardise and automate enterprise reporting.

    • Data Governance
    • PII
    • Ab Initio
    • Hive
  • Automation & ReliabilityAccenture

    AIX → Linux Migration Regression Utility

    An Ab Initio regression utility comparing AIX and Linux graph outputs and table loads, behind a zero-data-loss migration.

    data lost in the migration
    0
    • Ab Initio
    • Data Quality
    • AIX
    • Linux
  • AI / LLMSide project

    Local LLM Agent

    A side project: a tool-using agent built from scratch on a local LLM, including a portfolio-analysis tool that reads holdings from Excel.

    • Python
    • Ollama
    • LLM agents

04 / skills

Skills

primary skill

Big Data & Processing

  • Primary skill: PySpark
  • Apache Spark
  • Spark on YARN
  • Primary skill: Hive
  • HDFS
  • Alluxio
  • Hadoop

ETL & Data Integration

  • Primary skill: Ab Initio GDE
  • Express>It
  • EME
  • CDC (Oracle GoldenGate)
  • Solace Queue

Databases & Storage

  • Oracle
  • Aerospike
  • Netezza
  • Parquet

Orchestration & Ops

  • Apache Airflow
  • Primary skill: Unix/Linux shell scripting
  • cron

Cloud

  • GCP (reporting)
  • AWS S3

Languages

  • Primary skill: SQL
  • Python
  • PL/SQL

ML & AI Enablement

  • Feature-ready datasets
  • Data drift monitoring (KL, PSI, KS)
  • LLM agents
  • NL-to-SQL
  • Prompt engineering
  • Test harnesses for LLM apps

Data Management

  • Data quality
  • Data governance
  • PII encryption
  • Data migration & reconciliation
  • Performance tuning

Leadership

  • Team leadership & mentoring
  • Stakeholder reporting

05 / achievements

Achievements

  • India's first AI-powered spam detection

    Key contributor to a platform protecting 300M+ Airtel users.

  • Data drift framework

    Reusable framework adopted across multiple ML pipelines.

  • LLM data agent: 35/44 → 202/202 tests

    Built the automated test suite and hardened the agent to a full pass rate.

  • PII encryption across DARTS

    Implemented enterprise-wide PII protection.

  • Zero-data-loss AIX → Linux migration

    Plus the regression utility that verified it.

  • 2016

    Accenture Excellence Award

    Recognised for work at Accenture.

  • 2023–2024

    PGP in AI & ML, CGPA 3.79/4.0

    Great Lakes & The University of Texas at Austin.

06 / education

Education

  • 2023–2024

    Post Graduate Program in AI & Machine Learning

    Great Lakes Executive Learning & The University of Texas at Austin

    CGPA 3.79/4.0

  • 2010–2014

    B.E., Information Technology

    Sinhgad Academy of Engineering, University of Pune

07 / contact

Contact

Open to Lead / Senior Data Engineering roles — Bengaluru and other locations, on-site or hybrid.