001•THE OPENING
25 CASE STUDIES
SCENE 01 / 08|THE OPENING
CHENNAI / INDIA•3+ YEARS EXPERIENCE•BACKEND • DATA • CLOUD
SYSTEM POSITIONING // 2026

MUKESH KUMAR

SOFTWARE ENGINEER→DATA ENGINEER

3+ years of software engineering experience building production backend systems, relational schemas, and data pipelines — now specializing in distributed data processing, pipeline observability, and temporal incident replay.

EXPLORE DATA BLACK BOX→25 CASE STUDIES
STATUS: SYSTEM READY
SCROLL TO START JOURNEY
SCENE 02 / 08THE ENGINEER BEHIND THE SYSTEM

I DON'T JUST WRITE CODE.
I BUILD SYSTEMS THAT MOVE DATA,
SERVE APPLICATIONS, AND SCALE WITH USERS.

Mukesh Kumar Krishnamoorthi
MUKESH KUMAR
ENGINEER PORTRAIT
CAREER EVOLUTION

My software engineering journey began in production backend systems — crafting REST APIs, modeling relational databases, and debugging query bottlenecks.

As systems grew, I realized that modern software applications live and die by their data pipelines. I evolved into Data Engineering to solve distributed processing, pipeline reliability, data quality, and incident replay at scale.

FOUNDATIONBackend Engineering
SPECIALIZATIONData Platforms

CAREER PROGRESSION TIMELINE2023 — 2026

2023START

Software Engineering

Foundational software architecture, algorithm design, version control, and clean code practices.

2024CORE BACKEND

Backend Engineering & APIs

Node.js, TypeScript, RESTful services, PostgreSQL indexing, schema design, connection pooling.

2025ANALYTICS & ETL

Analytics & Data Transformation

Python Pandas batch jobs, SQL window aggregations, data validation layers, and automated reporting pipelines.

2026PRESENT SPECIALIZATION

Data Engineering & Observability

Apache SparkApache AirflowDatabricksAzure ADLSPostgreSQLData Black Box
SCENE 03 / 08TECHNICAL UNIVERSE

INTERACTIVE SYSTEMS MAP

Hover or tap any node to inspect deep engineering concepts, optimization strategies, and practical application details.

DATA ENGINEERING & PROCESSING
BACKEND & STORAGE
CLOUD & PLATFORM
DISTRIBUTED SYSTEMS & ARCHITECTURE
DATA TELEMETRY

Apache Spark

PRIMARY PURPOSE

Distributed processing & large-scale data transformations

ENGINEERING KNOWLEDGE & PATTERNS
Transformations vs Actions
Shuffle optimization
Partition tuning
Broadcast joins
Memory caching
Data skew handling
Spark UI profiling
STATUS: VERIFIED100% PRODUCTION READY
SCENE 04 / 08FLAGSHIP DATA SYSTEM
BLACKBOX FLIGHT RECORDER ACTIVE
PROJECT // BLACKBOX

DATA BLACK BOX

"GitHub for Data Incidents — replay exactly what happened inside a data pipeline."

Tagline: Every pipeline run leaves evidence. Replay the past. Fix the future.

PRODUCTION METRIC FAILURE DETECTEDRUN #8421 • 14:27 PM
2:13 PM: Revenue = ₹18.4M→2:27 PM: Revenue = ₹11.2M-39.1% DROP
FLIGHT METRICS LOGCOMMIT a91f3c
Started: 14:00:02
Completed: 14:07:31
Input Records: 2,841,992
Output Records: 2,113,401 (-728,591)
Spark Partitions: 199
Shuffle Read: 284 GB
Disk Spill: 31 GB
Data Quality Check: FAILED ❌
ROOT CAUSE IDENTIFIED BY DATA BLACK BOX
Column Modified:customer_status
EXPECTED IN INNER JOIN"active"
NEW UPSTREAM INGESTED VALUE"ACTIVE" (UPPERCASE)

Impact Chain: Ingested string case mutated from lower case to upper case. Inner join condition ON status = 'active' failed to match 728,591 records, dropping revenue aggregation by 39.1%.

CONFIDENCE SCORE: 97.4%
SCENE 05 / 08VERTICAL FILM REEL EXPERIENCE

ENGINEERING FILM TIMELINE

Real software engineering impact, ownership, and platform evolution from 2023 to present.

SOFTWARE ENGINEER (2023 — PRESENT)
ADYOG
STAGE 01

Production Software Foundations

Built robust web applications, structured frontend components, and validated client payload schemas.

STAGE 02

High-Throughput Backend APIs

Architected RESTful microservices in Node.js & TypeScript with modular routing and authentication filters.

STAGE 03

Database Systems & Query Tuning

Modeled relational schemas in PostgreSQL, optimized query execution plans, and tuned indexing for fast aggregations.

STAGE 04

Analytics & Data Transformation Pipelines

Engineered batch ETL workloads in Python Pandas and SQL, computing automated business analytics metrics.

STAGE 05 (PRESENT)

Distributed Data Engineering

Designing distributed processing jobs in PySpark, data pipeline orchestration in Airflow, and Data Black Box incident replay systems.

BACKEND IMPACT & OWNERSHIP
REST APIs
Node.js / TypeScript
PostgreSQL Schemas
Query Optimization
DATA PLATFORM IMPACT
ETL Pipelines
Python & Pandas
SQL Window Functions
Data Quality Assertions
SYSTEMS & DEVOPS OWNERSHIP
Docker Containers
Caching & Memory
Performance Profiling
System Design
SCENE 06 / 08THE ENGINEERING LAB

WHEN DATA GETS BIG

Interactive simulation demonstrating how Mukesh thinks as a Data Engineer — balancing shuffle, partitioning, broadcast joins, and data skew.

SCALE SIMULATOR // SPARK DISTRIBUTED ENGINE
STEP 1PARTITIONING2400 Partitions
STEP 2SHUFFLE REDUCTION1.1 TB Target
STEP 3BROADCAST JOINDimension Tables
STEP 4SKEW SALTINGResolve Hotspots
RESULTOPTIMIZED RUN9.0 min
UNOPTIMIZED SPARK RUN (1 TB)BEFORE
Shuffle Read: 4.8 TB
Execution Time: 27.0 min
Executors: 20
Partition Skew: High (Spill to Disk)
MUKESH OPTIMIZED SPARK RUN (1 TB)AFTER
Shuffle Read: 1.1 TB (-77%)
Execution Time: 9.0 min (3× Faster)
Executors: 20
Partition Strategy: Salting + AQE Enabled

Broadcast Strategy: Auto-broadcast hash join with predicate pushdown

Skew Mitigation: Salting join keys with random prefix (0-9) to resolve hotspot partitions

SCENE 08 / 08START A CONVERSATION

BUILD DATA SYSTEMS
THAT SCALE WITH YOUR VISION

Whether you need to scale distributed Spark pipelines, optimize database queries, or build reliable data platforms — let's connect.

DIRECT REACHOUT

EMAIL ADDRESSmukeshkumar.krishnamoorthi@gmail.com
GITHUB PROFILEgithub.com/mukeshkumar-krishnamoorthi
LINKEDIN PROFILElinkedin.com/in/mukesh-kumar-krishnamoorthi
LOCATIONChennai, India (Remote & Relocation)

SEND DIRECT TRANSMISSION

© 2026 MUKESH KUMAR. SOFTWARE ENGINEER → DATA ENGINEER.
BACK TO TOP ↑