irelink translates the Python splink
library into idiomatic R. This vignette maps common Splink 5 patterns to
irelink so you can get started quickly.
Splink uses an object-oriented design centered on a
Linker class. irelink uses a functional
pipeline that fits naturally in R. The Linker object’s
namespaced methods such as linker.training.* and
linker.inference.* become standalone functions that accept
and return an il_model object.
Splink bundles comparison levels into high-level comparison classes
such as JaroWinklerAtThresholds. In irelink,
the cl_*() functions fill the same role and can be passed
directly to il_compare().
| Step | splink (Python) | irelink (R) |
|---|---|---|
| Load data | splink_datasets.fake_1000 |
fake_1000 |
| Choose backend | DuckDBAPI() |
DBI::dbConnect(duckdb::duckdb()) |
| Register data | db_api.register(df, ...) |
handled by il_model() |
| Define settings | SettingsCreator(...) |
il_spec() |>il_compare(...) |>il_block_on(...) |
| Create model | Linker(df_sdf, settings) |
il_model(df, spec = spec, con = con) |
| Estimate prior | linker.training.estimate_probability_two_random_records_match(...) |
il_estimate_prior(model, ...) |
| Estimate u | linker.training.estimate_u_using_random_sampling(...) |
il_estimate_u(model) |
| Estimate m (EM) | linker.training.estimate_parameters_using_expectation_maximisation(...) |
il_estimate_em(model, ...) |
| Estimate m (labels) | linker.training.estimate_m_from_pairwise_labels(...) |
il_estimate_m_from_labels(model, ...) |
| Predict | linker.inference.predict(...) |
predict(model, ...) |
| Cluster | linker.clustering.cluster_pairwise_predictions_at_threshold(...) |
il_cluster(pairs) |
| Deterministic link | linker.deterministic_link() |
il_deterministic_link(df, ...) |
| Match new records | linker.inference.predict_between(df_sdf, new_sdf, ...) |
il_find_matches(model, new_records, ...) |
| Pairs within new records | linker.inference.predict_within(new_sdf, ...) |
il_attach(model, new_records) |>predict(...) |
| Score chosen pairs | linker.inference.score_pair(...),
score_pairs(...) |
il_score_pairs(model, records_l, records_r) |
Splink 5 requires registering each input with
db_api.register() before building a Linker,
and the Linker no longer takes db_api. In
irelink, il_model() registers data frames,
lazy tables, or table names on the supplied connection itself.
irelink also supports
link_type = "link_and_dedupe" for two-table jobs where
duplicates may exist within each input table and across the two
tables.
irelink scores in-memory inputs and DBI-backed tables,
including lazy DuckDB results. Splink 5’s chunked prediction, DuckDB
source pruning, and Parquet-backed intermediate tables are not
available, so very large workflows should rely on explicit blocking and
predict(collect = FALSE).
Comparison levels are the building blocks used to score how similar
two records are on a field. Each cl_*() function
corresponds to a Splink comparison level class.
| splink (Python) | irelink (R) |
|---|---|
ExactMatchLevel |
cl_exact() |
LevenshteinLevel |
cl_levenshtein() |
DamerauLevenshteinLevel |
cl_damerau_levenshtein() |
JaroLevel |
cl_jaro() |
JaroWinklerLevel |
cl_jaro_winkler() |
JaccardLevel |
cl_jaccard() |
CosineSimilarityLevel |
cl_cosine() |
AbsoluteDifferenceLevel |
cl_numeric_diff() |
PercentageDifferenceLevel |
cl_pct_diff() |
AbsoluteTimeDifferenceAtThresholds |
cl_date_diff() |
DistanceInKMLevel |
cl_geo_distance() |
ArrayIntersectLevel |
cl_array_intersect() |
CustomLevel |
cl_custom() |
NullLevel |
cl_null() |
ElseLevel |
cl_else() |
And |
cl_and() |
Or |
cl_or() |
Not |
cl_not() |
Splink provides high-level comparison classes for common field types.
In irelink, these are helper functions that return
preconfigured sets of levels.
| splink (Python) | irelink (R) |
|---|---|
NameComparison |
cl_name() |
ForenameSurnameComparison |
cl_forename_surname() |
DateOfBirthComparison |
cl_dob() (Levenshtein for one-character typos) |
EmailComparison |
cl_email() |
PostcodeComparison |
cl_postcode() |
| splink (Python) | irelink (R) |
|---|---|
linker.visualisations.match_weights_chart() |
il_weights(model) |
linker.visualisations.parameter_estimate_comparisons_chart() |
il_parameters(model) |
linker.visualisations.waterfall_chart(...) |
il_waterfall(pairs, ...) |
| comparison levels for a pair, without a model | il_compare_records(record_a, record_b, spec) |
linker.evaluation.prediction_errors_from_labels_column(...) |
il_errors(model, ...) |
linker.evaluation.unlinkables_chart() |
il_unlinkables(model) |
Splink 5 combines these analyses in
linker.evaluation.accuracy_analysis_from_labels_column(),
selected with output_type.
| splink (Python) | irelink (R) |
|---|---|
accuracy_analysis_from_labels_column(..., output_type="accuracy") |
il_accuracy(model, ...) |
accuracy_analysis_from_labels_column(..., output_type="precision_recall") |
il_precision_recall(model, ...) |
accuracy_analysis_from_labels_column(..., output_type="roc") |
il_roc(model, ...) |
| splink (Python) | irelink (R) |
|---|---|
splink.exploratory.profile_columns(...) |
il_profile(df, ...) |
splink.exploratory.completeness_chart(...) |
il_completeness(df, ...) |
splink.blocking_analysis.count_comparisons_from_blocking_rules(...) |
il_count_pairs(df, ...) |
splink.blocking_analysis.n_largest_blocks(...) |
il_largest_blocks(df, ...) |
Splink 5 estimates blocking comparison counts from a 5% record sample
by default. il_count_pairs() computes exact counts unless
you set record_sample_proportion below 1.
| splink (Python) | irelink (R) |
|---|---|
linker.misc.save_model_to_json(...) |
il_save(model, path) |
Linker(df_sdf, "model.json") |
il_load(path) |
linker.table_management.delete_tables_created_by_splink_from_db() |
il_cleanup_all(con) |
| model-scoped cleanup | il_cleanup(model) |
In Splink, you create blocking rules with block_on(),
and irelink uses the same function name. The main
difference is where the rules are used: Splink passes them into
SettingsCreator, while irelink adds them to a
spec with il_block_on() or passes them directly to training
functions.
Below is a minimal deduplication example in both Splink and
irelink.
splink (Python):
from splink import Linker, SettingsCreator, DuckDBAPI, block_on, splink_datasets
import splink.comparison_library as cl
db_api = DuckDBAPI()
df_sdf = db_api.register(splink_datasets.fake_1000, dataset_display_name="fake_1000")
settings = SettingsCreator(
link_type="dedupe_only",
comparisons=[
cl.JaroWinklerAtThresholds("first_name", [0.9, 0.7]),
cl.JaroWinklerAtThresholds("surname", [0.9, 0.7]),
cl.ExactMatch("dob"),
],
blocking_rules_to_generate_predictions=[
block_on("first_name"),
block_on("surname"),
],
)
linker = Linker(df_sdf, settings)
linker.training.estimate_u_using_random_sampling(max_pairs=1e6)
linker.training.estimate_parameters_using_expectation_maximisation(
block_on("surname")
)
pairwise = linker.inference.predict(threshold_match_probability=0.5)
clusters = linker.clustering.cluster_pairwise_predictions_at_threshold(
pairwise, 0.95
)irelink (R):
library(irelink)
df <- fake_1000
con <- DBI::dbConnect(duckdb::duckdb())
spec <- il_spec() |>
il_compare(first_name, cl_jaro_winkler(0.9, 0.7)) |>
il_compare(surname, cl_jaro_winkler(0.9, 0.7)) |>
il_compare(dob, cl_exact()) |>
il_block_on(first_name) |>
il_block_on(surname)
model <- il_model(df, spec = spec, con = con)
model <- il_estimate_u(model)
model <- il_estimate_em(model, block_on(surname))
pairs <- predict(model, threshold = 0.5)
clusters <- il_cluster(pairs)
il_cleanup(model)
DBI::dbDisconnect(con, shutdown = TRUE)The examples above use probability thresholds because those transfer
cleanly between Splink and irelink. In Splink, prediction
match_weight includes the prior odds. In
irelink, match_weight is evidence only, and
total_match_weight is the prior-inclusive log2 odds. Keep
that difference in mind if you translate match-weight thresholds between
the two packages.
splink (Python):
new_sdf = db_api.register(
[{"unique_id": 1001, "first_name": "Jhon", "surname": "Smith", "dob": "1990-01-15"}],
dataset_display_name="new_records",
)
results = linker.inference.predict_between(
df_sdf, new_sdf, threshold_match_probability=0.5
)irelink (R):
new_df <- data.frame(
first_name = "Jhon",
surname = "Smith",
dob = "1990-01-15"
)
results <- il_find_matches(model, new_df, threshold = 0.5)il_find_matches() corresponds to
predict_between(): it scores new records against the
model’s existing data, but not new records against each other. For
Splink 5’s predict_within(), attach the new records to the
trained model with il_attach() and call
predict(). Note that il_attach() computes term
frequencies from the attached records. When those are too few to be
representative, replace them with term frequencies from the full data
using il_register_tf(..., overwrite = TRUE).