splink
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
File Explorer
Download Latest Version (.zip)- bug_report.yml
- config.yml
- FEATURE_REQUEST.md
- check-lockfile.yml
- check-minimal-import.yml
- codeql-analysis.yml
- demo-examples.yml
- demo-tutorials.yml
- dependency-canary.yml
- dependency-review.yml
- documentation.yml
- lint.yml
- pypi-release.yml
- pytest-duckdb.yml
- pytest-postgres.yml
- pytest-spark.yml
- pytest-sqlite.yml
- typechecking.yml
- dependabot.yml
- PULL_REQUEST_TEMPLATE.md
- api_docs_index.md
- blocking.md
- blocking_analysis.md
- clustering.md
- column_expression.md
- comparison_level_library.md
- comparison_library.md
- datasets.md
- em_training_session.md
- evaluation.md
- exploratory.md
- inference.md
- linker_clustering.md
- misc.md
- settings_dict_guide.md
- splink_dataframe.md
- table_management.md
- training.md
- visualisations.md
- bias_chart.png
- bias_investigation_steps.png
- bias_mitigation_flowchart.png
- charts_gallery.png
- completeness_chart.png
- confusion_matrix.png
- confusion_matrix_1.png
- confusion_matrix_2.png
- data-linking-pipeline.png
- justice-user-journey1.drawio.png
- linkage_process.drawio.png
- linking_complexity_overview.png
- match_weights_chart.png
- moj-linking-tech-stack.excalidraw.png
- Person_Timeline.drawio.png
- process_flow.png
- record_eg.png
- sp_mistake_data.png
- transitive_links.drawio.png
- 2023-07-27-feature_update.md
- 2023-12-06-feature_update.md
- 2024-01-25-ethics.md
- 2024-03-19-splink4.md
- 2024-07-10-splink4_release.md
- 2024-07-11-bias.md
- 2024-08-15-bias-continued.md
- 2026-01-29-running-splink-in-production.md
- 2026-06-17-splink-5-release.md
- .authors.yml
- index.md
- accuracy_chart_from_labels_table.png
- cluster_studio.html
- cluster_studio_dashboard.png
- comparator_score_chart.png
- comparator_score_threshold_chart.png
- comparison_viewer_dashboard.png
- completeness_chart.png
- completeness_chart_tooltip.png
- cumulative_num_comparisons_from_blocking_rules_chart.png
- cumulative_num_comparisons_from_blocking_rules_chart_tooltip.png
- m_u_parameters_chart.png
- m_u_parameters_chart_tooltip_1.png
- m_u_parameters_chart_tooltip_2.png
- match_weights_chart.png
- match_weights_chart_tooltip.png
- missingness_chart.png
- missingness_chart_tooltip.png
- parameter_estimate_comparisons_chart.png
- phonetic_match_chart.png
- profile_columns.png
- profile_columns_tooltip_1.png
- profile_columns_tooltip_2.png
- profile_columns_tooltip_3.png
- roc_chart_from_labels_table.png
- roc_chart_from_labels_table_tooltip.png
- roc_curve_explainer.png
- scv.html
- tf_adjustment_chart.png
- tf_adjustment_chart_tooltip_1.png
- tf_adjustment_chart_tooltip_2.png
- threshold_selection_tool_from_labels_table.png
- unlinkables_chart.png
- unlinkables_chart_tooltip.png
- waterfall_chart.png
- waterfall_chart_tooltip.png
- accuracy_analysis_from_labels_table.nb.py
- cluster_studio_dashboard.nb.py
- comparison_viewer_dashboard.nb.py
- completeness_chart.nb.py
- cumulative_comparisons_to_be_scored_from_blocking_rules_chart.nb.py
- index.md
- m_u_parameters_chart.nb.py
- match_weights_chart.nb.py
- parameter_estimate_comparisons_chart.nb.py
- profile_columns.nb.py
- tf_adjustment_chart.nb.py
- threshold_selection_tool_from_labels_table.nb.py
- unlinkables_chart.nb.py
- waterfall_chart.nb.py
- custom.css
- neoteroi-mkdocs.css
- .gitignore
- source.txt
- fake_1000_combined.json
- model_create_h50k.nb.py
- model_h50k.json
- real_time_settings.json
- saved_model_from_demo.json
- 50k_cluster.html
- 50k_deterministic_cluster.html
- comparison_viewer_transactions.html
- accuracy_analysis_from_labels_column.nb.py
- deduplicate_50k_synthetic.nb.py
- deterministic_dedupe.nb.py
- febrl3.nb.py
- febrl4.nb.py
- link_only.nb.py
- pairwise_labels.nb.py
- quick_and_dirty_persons.nb.py
- real_time_record_linkage.nb.py
- transactions.nb.py
- pseudopeople_cluster_studio.html
- pseudopeople_scv.html
- bias_eval.nb.py
- business_rates_match.nb.py
- comparison_playground.nb.py
- cookbook.nb.py
- pseudopeople-acs.nb.py
- deduplicate_1k_synthetic.nb.py
- 50k_cluster.html
- deduplicate_50k_synthetic.nb.py
- examples_index.md
- 00_Tutorial_Introduction.nb.py
- 01_Prerequisites.nb.py
- 02_Exploratory_analysis.nb.py
- 03_Blocking.nb.py
- 04_Estimating_model_parameters.nb.py
- 05_Predicting_results.nb.py
- 06_Visualising_predictions.nb.py
- 07_Evaluation.nb.py
- 08_building_your_own_model.md
- 09_scaling_up_techniques.nb.py
- cluster_studio.html
- scv.html
- conftest.py
- pyproject.toml
- uv.lock
- blog_posts.md
- contributing_to_docs.md
- development_quickstart.md
- lint_and_format.md
- managing_dependencies_with_uv.md
- releases.md
- testing.md
- building_charts.nb.py
- understanding_and_editing_charts.md
- caching.md
- CONTRIBUTING.md
- debug_modes.md
- dependency_compatibility_policy.md
- index.md
- spark_pipelining_and_caching.md
- transpilation.md
- udfs.md
- __init__.py
- cumulative_comparisons.png
- pairwise_comparisons.png
- AltairUserGuide.png
- chart.json
- charts.mp4
- new_chart.png
- new_chart_def.json
- old_chart.png
- old_chart_def.json
- Vega-Lite-editor.png
- basic_graph.drawio.png
- basic_graph_centralisataion.drawio.png
- basic_graph_cluster.drawio.png
- basic_graph_cluster_person.drawio.png
- basic_graph_records.drawio.png
- cluster_density.drawio.png
- cluster_size.drawio.png
- graph.png
- is_bridge.drawio.png
- threshold_cluster.drawio.png
- threshold_cluster_high.drawio.png
- threshold_cluster_low.drawio.png
- threshold_cluster_medium.drawio.png
- python_release_cycle.png
- prob_v_weight.png
- waterfall.png
- probabilistic_example.png
- simplified_waterfall.png
- splink_01_input_records.png
- splink_02_pairwise_links.png
- splink_03_clusters.png
- notes.png
- notes_button.png
- publish.png
- tag.png
- error_logger.png
- calc.png
- example.png
- gender-distribution.png
- surname-distribution.png
- tf-intro.drawio.png
- tf-match-weight.png
- waterfall.png
- favicon.ico
- postcode_components.png
- vega_spec_for_readme.vg.json
- .gitignore
- dataset_labels_table.md
- datasets_table.md
- tags.md
- mathjax.js
- main.html
- blocking_rules.md
- model_training.md
- performance.md
- choosing_comparators.nb.py
- comparators.md
- comparisons_and_comparison_levels.md
- customising_comparisons.nb.py
- out_of_the_box_comparisons.nb.py
- phonetic.md
- regular_expressions.nb.py
- term-frequency.md
- feature_engineering.md
- graph_metrics.md
- how_to_compute_metrics.nb.py
- overview.md
- confusion_matrix.drawio.png
- confusion_matrix_extra.drawio.png
- edge_metrics.md
- edge_overview.md
- labelling.md
- model.md
- overview.md
- drivers_of_performance.md
- optimising_duckdb.md
- optimising_spark.md
- performance_of_comparison_functions.nb.py
- backends.md
- postgres.md
- link_type.md
- querying_splink_results.md
- settings.md
- fellegi_sunter.md
- linked_data_as_graphs.md
- probabilistic_vs_deterministic.md
- record_linkage.md
- training_rationale.md
- topic_guides_index.md
- getting_started.md
- index.md
- profile_columns_tooltip_1.png
- development_environment.yaml
- development_environment_lock_Linux-x86_64.txt
- development_setup_with_conda.sh
- test_uv_build.sh
- docker-compose.yaml
- servers.json
- setup.sh
- teardown.sh
- custom_dictionary.txt
- pyspelling.yml
- spellchecker.sh
- ensure_packages_installed.sh
- build-test.dockerfile
- generate_dataset_docs.py
- make_docs_locally.sh
- make_test_datasets_smaller.py
- reduce_notebook_runtime.py
- run_notebooks.sh
- test_build.sh
- update_vega.py
- duckdb.py
- postgres.py
- spark.py
- sqlite.py
- __init__.py
- enable_splink.py
- __init__.py
- metadata.py
- splink_datasets.py
- utils.py
- __init__.py
- duckdb_helpers.py
- __init__.py
- database_api.py
- database_api_with_profiling.py
- dataframe.py
- .gitignore
- accuracy_chart.json
- blocking_rule_generated_comparisons.json
- comparator_score_chart.json
- comparator_score_threshold_chart.json
- completeness.json
- m_u_parameters_interactive_history.json
- match_weight_histogram.json
- match_weights_interactive_history.json
- match_weights_waterfall.json
- parameter_estimate_comparisons.json
- phonetic_match_chart.json
- precision_recall.json
- probability_two_random_records_match_iteration.json
- profile_data.json
- profile_data_outer.json
- roc.json
- tf_adjustment_chart.json
- threshold_selection_tool.json
- unlinkables_chart_def.json
- d3@7.8.5
- stdlib.js@5.8.3
- vega-embed@6.20.2
- vega-lite@5.2.0
- vega@5.31.0
- slt.js
- template.j2
- scala-udf-similarity-0.1.2_spark3.x.jar
- scala-udf-similarity-0.2.0_spark4.x.jar
- scala-udf-similarity-0.2.1_spark4.x.jar
- cluster_template.j2
- custom.css
- custom.css
- template.j2
- splink_vis_utils.js
- single_chart_template.html
- DEPENDENCY_LICENSES.txt
- __init__.py
- blocking_analysis.py
- clustering.py
- evaluation.py
- inference.py
- misc.py
- table_management.py
- training.py
- visualisations.py
- __init__.py
- database_api.py
- dataframe.py
- __init__.py
- log_invalid_columns.py
- settings_column_cleaner.py
- settings_validation_log_strings.py
- valid_types.py
- __init__.py
- custom_spark_dialect.py
- version.py
- __init__.py
- database_api.py
- database_api_with_profiling.py
- dataframe.py
- jar_location.py
- __init__.py
- database_api.py
- dataframe.py
- __init__.py
- accuracy.py
- block_from_labels.py
- blocking.py
- blocking_analysis.py
- blocking_rule_creator.py
- blocking_rule_creator_utils.py
- blocking_rule_library.py
- cache_dict_with_logging.py
- charts.py
- chunking.py
- cluster_studio.py
- clustering.py
- column_expression.py
- comparison.py
- comparison_creator.py
- comparison_level.py
- comparison_level_composition.py
- comparison_level_creator.py
- comparison_level_library.py
- comparison_level_sql.py
- comparison_library.py
- comparison_vector_distribution.py
- comparison_vector_values.py
- completeness.py
- connected_components.py
- constants.py
- database_api.py
- dialects.py
- edge_metrics.py
- em_sampling.py
- em_training_session.py
- estimate_u.py
- exceptions.py
- expectation_maximisation.py
- find_matches_to_new_records.py
- graph_metrics.py
- input_column.py
- labelling_tool.py
- linker.py
- logging_messages.py
- lower_id_on_lhs.py
- m_from_labels.py
- m_training.py
- m_u_records_to_parameters.py
- match_weights_histogram.py
- misc.py
- one_to_one_clustering.py
- parse_sql.py
- pipeline.py
- predict.py
- profile_data.py
- realtime.py
- settings.py
- settings_creator.py
- similarity_analysis.py
- splink_comparison_viewer.py
- splink_dataframe.py
- splink_logging.py
- splinkdataframe_utils.py
- sql_transform.py
- term_frequencies.py
- testing.py
- unique_id_concat.py
- unlinkables.py
- vertically_concatenate.py
- waterfall_chart.py
- __init__.py
- blocking_analysis.py
- blocking_rule_library.py
- clustering.py
- comparison_level_library.py
- comparison_library.py
- datasets.py
- exploratory.py
- logging.py
- py.typed
- postgres_conf.py
- arrays_df.parquet
- fake_1000_from_splink_demos.csv
- fake_1000_from_splink_demos_strip_datetypes.parquet
- known_params_comparison_vectors.csv
- splink2_479_vs_481.csv
- splink2_m_u_history_fixed_u.csv
- splink2_m_u_history_no_fix.csv
- splink2_proportion_of_matches_history_fixed_u.csv
- splink2_proportion_of_matches_history_no_fix.csv
- test_array.parquet
- __init__.py
- basic_settings.py
- cc_testing_utils.py
- conftest.py
- decorator.py
- helpers.py
- literal_utils.py
- test_accuracy.py
- test_analyse_blocking.py
- test_array_based_blocking.py
- test_array_columns.py
- test_blocking.py
- test_blocking_rule_composition.py
- test_caching.py
- test_cc_random_graphs.py
- test_charts.py
- test_chunking.py
- test_cluster_at_multiple_thresholds.py
- test_cluster_studio.py
- test_cluster_using_single_best_links.py
- test_clustering.py
- test_column_expression.py
- test_columns_selected.py
- test_columns_used.py
- test_compare_splink2.py
- test_comparison_level.py
- test_comparison_level_composition.py
- test_comparison_level_lib.py
- test_comparison_lib.py
- test_comparison_template_lib.py
- test_comparison_viewer_dashboard.py
- test_completeness.py
- test_compound_comparison_levels.py
- test_correctness_of_convergence.py
- test_dataframe_in_out_formats.py
- test_date_levels_and_comparisons.py
- test_debug_mode.py
- test_disable_tf_exact_match_detection.py
- test_em_max_pairs.py
- test_estimate_prob_two_rr_match.py
- test_expectation_maximisation.py
- test_extreme_match_weights.py
- test_full_example_deterministic_link.py
- test_full_example_duckdb.py
- test_full_example_postgres.py
- test_full_example_spark.py
- test_full_example_sqlite.py
- test_graph_metrics.py
- test_input_column.py
- test_join_type_for_estimate_u_and_predict_are_efficient.py
- test_km_distance_level.py
- test_lat_long_distance.py
- test_linker_variants.py
- test_m_train.py
- test_new_comparison_levels.py
- test_new_db_api.py
- test_postgres_udfs.py
- test_predict_between.py
- test_predict_within.py
- test_profile_data.py
- test_realtime.py
- test_regex_param.py
- test_score_missing_edges.py
- test_score_pairs.py
- test_settings_options.py
- test_settings_validation.py
- test_spark_udfs.py
- test_splink_datasets.py
- test_sql_transform.py
- test_term_frequencies.py
- test_testing_fns.py
- test_total_comparison_count.py
- test_train_vs_predict.py
- test_u_train.py
- utils.py
- .dockerignore
- .gitignore
- CHANGELOG.md
- CONTRIBUTING.md
- LICENSE
- Makefile
- mkdocs.yml
- pyproject.toml
- README.md
- uv.lock
# Installation Guide
git clone https://github.com/moj-analytical-services/splink
Downloads the entire project code from GitHub to your computer.
cd splink
Moves into the project folder you just downloaded.
2. Official Install Script
Easy Recommended- Python 3 Python is required to use pip.
pip install splink
Installs the package published on PyPI directly β no need to clone the source.
pip install 'splink[{backend}]'
Installs the package published on PyPI directly β no need to clone the source.
pip install 'splink[spark]'
Installs the package published on PyPI directly β no need to clone the source.
pip install 'splink[postgres]'
Installs the package published on PyPI directly β no need to clone the source.
Pulled directly from this repo's README.
3. Docker
Easy- Git Needed to download the project code from GitHub.
- Docker Desktop Needed to build and run containers. Install it and keep it running in the background.
docker compose -f scripts/postgres_docker/docker-compose.yaml up -d --build
Runs the command against the services defined in the compose file.
4. Python
Easypip install splink
Installs the package published on PyPI directly β no need to clone the source.
pip install 'splink[{backend}]'
Installs the package published on PyPI directly β no need to clone the source.
pip install 'splink[spark]'
Installs the package published on PyPI directly β no need to clone the source.
pip install 'splink[postgres]'
Installs the package published on PyPI directly β no need to clone the source.
Pulled directly from this repo's README.
5. Make
Medium- Git Needed to download the project code from GitHub.
- Make Usually pre-installed on Linux/macOS. On Windows, install separately (e.g. via MSYS2 or WSL).
make
Compiles the code based on the generated build configuration to produce an executable.
