30% faster pipelines & zero bottlenecks: data platform modernization at Morgan Stanley
Morgan Stanley had 10+ legacy systems feeding disconnected pipelines at petabyte scale – and two global engineering squads across Mexico and India with no clear delivery framework. The mandate was to fix both simultaneously, without pausing delivery to VP-level stakeholders across investment banking and mobile app operations.
Morgan Stanley's enterprise data layer consisted of fragmented, poorly documented pipelines integrating 10+ heterogeneous source systems at petabyte scale. Batch jobs ran on aging Python 2.x with no unified transformation standard. Teams split between Mexico and India lacked clear delivery frameworks, creating bottlenecks at every VP stakeholder touchpoint. The organization needed both a platform overhaul and a delivery model transformation – simultaneously.
Led the end-to-end platform redesign on Azure Databricks implementing Medallion Architecture (Bronze / Silver / Gold) using PySpark and Delta Lake. Designed metadata-driven source-to-target mapping logic across all 10+ heterogeneous systems; applied complex SQL optimization (partitioning, clustering, indexing) at petabyte scale. Built ETL/ELT pipelines with Talend, dbt, and Airflow. Simultaneously served as Service Delivery Manager and Scrum Master: designed a requirements flowchart that eliminated bottlenecks across 6+ VP stakeholders, led the Python 2 to 3 migration, and drove squad alignment through SAFe ceremonies.
The platform redesign delivered a measurable 30% improvement in pipeline efficiency. The metadata-driven mapping approach standardized data integration across all 10+ systems, cutting new source onboarding time by weeks. The revamped delivery model eliminated stakeholder bottlenecks and increased squad throughput. The Snowflake analytics layer enabled executive dashboards for VP-level decision-making with near-real-time data. The Python 2 to 3 migration eliminated all end-of-life security vulnerabilities across critical systems.