Data Engineer, Ernst & Young (EY) – Client: Morgan Stanley
Present
- Consolidated heterogeneous flat files and databases into the enterprise data warehouse to ensure reliable downstream access and data availability.
- Led a team of six members, gathering comprehensive requirements and ensuring the consistent, timely delivery of project milestones.
- Developed and optimized PySpark-based data processing workflows on Databricks for data transformation, validation, and reliable delivery of curated datasets to downstream systems.
- Supported Databricks-based data migration and integration activities, applying ETL and data warehousing patterns to process, validate, and prepare large-scale datasets for downstream consumption.
- Spearheaded large-scale data migrations from Teradata and Hadoop to Snowflake, handing over 300 tables and 1 TB of data with minimal downtime and data inconsistency post-cut over.
- Built a unified application using ChatGPT and Claude AI agents to capture data lineage, expanding coverage from 1 to 10 Product teams.
- Implemented Snowflake LLMs and a Cortex agent to automate cross-system data lineage, improving traceability and query ability.
- On boarded new product data end-to-end into the warehouse, ensuring accurate schema mapping and timely data availability.
- Developed and optimized ETL/ELT pipelines using Informatica and PySpark, improving processing reliability and enabling on-time delivery of critical downstream data extracts.
- Optimized the legacy job framework to remove redundant steps and cut batch processing time from 10 hours to 7 hours.