30% faster pipelines & zero bottlenecks: data platform modernization at Morgan Stanley
Morgan Stanley had 10+ legacy systems feeding disconnected pipelines at petabyte scale – and two global engineering squads across Mexico and India with no clear delivery framework. The mandate was to fix both simultaneously, without pausing delivery to VP-level stakeholders across investment banking and phone application operations.
Morgan Stanley's enterprise data layer consisted of fragmented, poorly documented pipelines integrating 10+ heterogeneous source systems at petabyte scale. Batch jobs ran on aging Python 2.x with no unified transformation standard. Cross-globe teams in Mexico and India lacked clear delivery frameworks, creating bottlenecks at every VP stakeholder touchpoint. The organization needed both a platform overhaul and a delivery model transformation – simultaneously.
Led the end-to-end platform redesign on Azure Databricks implementing Medallion Architecture (Bronze / Silver / Gold) using PySpark and Delta Lake. Designed metadata-driven source-to-target mapping logic across all 10+ heterogeneous systems; applied complex SQL optimization (partitioning, clustering, indexing) at petabyte scale. Built ETL/ELT pipelines with Talend, dbt, and Airflow. Simultaneously served as Service Delivery Manager and Scrum Master: designed a requirements flowchart that eliminated bottlenecks across 6+ VP stakeholders, led the Python 2 to 3 migration, and drove squad alignment through SAFe ceremonies.
The platform redesign delivered a measurable 30% improvement in pipeline efficiency. The metadata-driven mapping approach standardized data integration across all 10+ systems, cutting new source onboarding time by weeks. The revamped delivery model eliminated stakeholder bottlenecks and increased squad throughput. The Snowflake analytics layer enabled executive dashboards for VP-level decision-making with near-real-time data. The Python 2 to 3 migration eliminated all end-of-life security vulnerabilities across critical systems.