AI-Powered Entity Resolution & Production ML Pipelines
Innovare's InnoTM application matched organizations and contacts across multiple data sources – and the rule-based system was quietly failing. Mismatches were corrupting downstream reports and nobody knew how deep the problem went. They needed matching that improved iteratively, with full production pipelines and automated data quality enforcement.
Innovare's InnoTM application needed to resolve and match entities (organizations, contacts, records) across multiple heterogeneous data sources with varying formats, naming conventions, and completeness levels. Rule-based matching was failing due to data inconsistencies. The team needed model-driven matching that could improve accuracy iteratively, alongside reliable production pipelines and automated data quality gates – all deployed to production on a tight timeline.
Designed and deployed algorithmic and model-driven matching solutions: implemented similarity scoring, probabilistic matching, and clustering at scale using Python (NumPy, Pandas, SciPy, Scikit-learn) and advanced SQL. Developed semantic data models and reusable transformation patterns to improve matching quality and consistency. Built production-grade pipelines end-to-end, from requirements and architecture through coding, testing, deployment, and orchestration with Apache Airflow. Introduced automated data quality testing to enforce integrity across all transformation outputs.
Model-based entity matching replaced failing rule-based systems with iteratively improvable models that outperformed previous accuracy benchmarks. Production pipelines ran end-to-end with full Airflow orchestration, automated testing, and documented data contracts. The semantic modelling and reusable transformation patterns reduced the cost of adding new data sources from weeks to days. The team established shared best practices promoted across the engineering chapter.