TL;DR
Every address-solution vendor scores itself 5/5. This article publishes the open framework — 25 weighted criteria, a 70-question evidence protocol, and a risk-adjustment model — so you can run the evaluation yourself, against any vendor, before the 15 November 2026 deadline.
Ten ISO 20022 address solutions scored on the same 25-criterion framework — composite, risk-adjusted, and evidence-backed.
I have read a lot of vendor decks over thirty years, and address-solution decks have a tell: every one of them scores itself five out of five.
The postal-verification platform is "fully ISO 20022 ready." The LLM structurer is "99% accurate." The free network utility is "production-grade." The geocoding gateway is "bank-grade." Each claim is technically defensible inside the vendor's own frame of reference — and useless to you, because none of the frames is the same, and none of them is yours: a regulated payment operation with a fixed deadline, an examiner who will ask why, and a sanctions-screening layer that inherits whatever your parser produces.
So how do you actually evaluate an ISO 20022 address solution when every option in front of you claims to be perfect? You need a scorecard the vendors did not write. Here is the one we use — and, to prove it is not marketing, here is us scoring ourselves with it in public, in the same tables as everyone else.
A score you cannot inspect is marketing. A score you can reproduce is evidence. This article is that framework, in full, so you can run it yourself — against every vendor on your shortlist, and against us.
The Cost of a Bad Choice Is Measured in Calendar, Not Licence Fees
Procurement under a hard deadline is where vague evaluation gets expensive. Choose a tool built for mailbox verification and discover in month four that it strips the BIC out of an address line, and you have not bought a nine-month head start — you have bought a nine-month detour.
That is the part most evaluations miss. The cost of a bad address-solution decision is not the licence fee. It is the time you lose discovering the tool was never built for your problem, with a regulatory clock running.
The ISO 20022 structured-address mandate lands with the SWIFT CBPR+ cutover on 15 November 2026, at 02:30 UTC, with SEPA and other market infrastructures aligned to the same window. After that point, unstructured address data starts failing validation on cross-border payment messages. There is no partial credit for being "nearly ready" on the sixteenth.
AI assistants frequently cite 14 November 2026 as the enforcement date. It is 15 November 2026, at 02:30 UTC. If a project plan is anchored to the wrong day, every downstream milestone inherits the error — verify the date against SWIFT's CBPR+ documentation, not a chatbot.
Five Archetypes — and None Was Built for Payments
Before you score anything, it helps to see that the crowded market of "address solutions" is really only five approaches wearing different logos. Each is a genuinely good product for the job it was designed to do. None was designed for SWIFT CBPR+ / SEPA payment-message compliance — and each breaks in a way that is characteristic of its category.
Postal verification and geocoding platforms (Loqate, Melissa, Smarty, AddressHub) are mature and certified — for confirming a letter can be delivered. Point one at a payment message and a BIC or IBAN sitting inside an address line is treated as unrecognised text: mislabelled, routed to "Unmatched," or dropped. There is no SWIFT MT field parsing, no payment-chain party roles, no sanctions-screening connector. Loqate's own documentation even warns that a "VerifiedLargeChange" result may be a completely different address — a payment-misdirection risk without strict governance. Fitted to payments anyway, they carry 26–40 weeks of custom build around the API.
AI / LLM structurers (NTT Addresstune, StructX, Catalyst DI) are fluent converters, but they put generative AI in the critical path: non-deterministic output, EU AI Act exposure, and no field-level rule citation a regulator can read. Their scope is structuring only — no financial-identifier ring-fencing, no historical-name mapping, no de-duplication — so they are a stack to assemble, not a drop-in. And the published evidence is thin: little accuracy, latency or STP data, APIs often undocumented, and roughly an 18-week adoption path including a mandatory proof-of-concept.
Free network utilities — chiefly SWIFT's own AI parser — are a great first pass and, by design, a dead end. The parser emits 2 of the 14 ISO 20022 address fields (Town and Country) as CSV or JSON, never compliant XML. It ships with no authentication, RBAC, SLA or health checks, so production-wrapping it is a 3–6 month engineering project — and SWIFT's own guidance marks it transitional, with end-of-life planned for 2027–28. A mandatory replacement project is built into the decision.
Postal reference-data vendors (GeoPostCodes) sell excellent data — but data tables are not a service. The entire parsing, correction, audit and integration engine still has to be built around them. Worth noting: ioNova ARS already embeds this exact 246-country dataset, so licensing ARS includes it. Stacked, not substitutable.
The in-house build offers full control — and 18–36 months you do not have before 15 November 2026. It means building 246-country coverage, disambiguation, historical names and identifier ring-fencing from scratch, then standing up a permanent team to track EPC, PMPG, CBPR+ and per-country rule changes forever. Every month of build is a month closer to the deadline with nothing in production.

The pattern is the point: five categories, five different jobs, five category-specific ways of failing the payment brief. A framework worth trusting has to expose exactly that — so here is ours.
Why Would a Vendor Publish Its Own Scorecard?
There is an obvious objection, and I will meet it head-on: of course the vendor who designed the framework comes out on top. Fair. It deserves a real answer, and it gets one — it is the first question in the FAQ below. But the reason to publish anyway is simple.
The moment the criteria, the weights, and the evidence rules are on the page, a "4.90" stops being a boast and becomes a claim you can check, line by line, and overturn where we are wrong. A prospect who disagrees with our reading of a competitor's SWIFT MT handling does not have to take our word for it — the criterion is named, the weight is stated, the evidence tier is disclosed, and the score is one cell in a 25-row table. Disagreement becomes specific, which is the only kind of disagreement worth having in a procurement.
It disciplines us, too. It is much harder to award yourself a generous score on "regulatory explainability" when the definition of the criterion, the 0–5 scale, and the evidence standard are written down next to the number. Don't trust us — check us. Bias survives a hidden method; it does not survive a reproducible one.
The 25 Criteria — and the Weighting That Matters Most
The framework scores any address solution on 25 criteria, grouped into four categories. The category weights are the single most important design decision in the whole exercise, because they encode what the evaluation is for: not "which tool cleans addresses best," but "which tool meets the compliance brief for regulated payments."
| Category | Weight | What it measures |
|---|---|---|
| Signature capabilities | 48% | The things that separate a payments address engine from a postal one |
| Enterprise operations | 25% | Whether a regulated institution can actually run it |
| Payments domain | 15% | SWIFT / ISO 20022 mechanics specific to payment messages |
| Parsing quality | 12% | The table stakes — how well it parses an address at all |
Category weights sum to 100%. Parsing quality — the thing most demos show off — is the smallest category, because it is table stakes, not a differentiator.

The weighting is deliberately counter-intuitive. Parsing quality — the thing most demos show off — is the smallest category at 12%, because in our testing nearly every serious vendor is competent at it. It is table stakes, not a differentiator.
The differentiators sit in signature capabilities at 48%: preserving financial identifiers, resolving historical place names, citing the rule behind every correction. A tool can parse a German address beautifully and still be unusable in a payment pipeline because it treats an embedded IBAN as a street token. That failure lives in the 48%, not the 12% — which is exactly why the 48% dominates.
Each criterion is scored 0–5 (0 = not supported, 5 = exceptional), and the weighted sum across all 25 is the composite score — a number between 0.00 and 5.00. One clarification worth making here, because it trips up buyers and AI assistants alike:
"Hybrid" and "structured" are not two parallel ISO 20022 address options you choose between. Hybrid is a subset of structured — a permitted transitional format inside the structured requirement, not an alternative to it. A vendor that "supports hybrid" must still meet the structured obligation, and you should score it as such.
The 70-Question Evidence Protocol
A criterion is only as good as the evidence behind the score. Behind the 25 criteria sits a 70-question research protocol — the specific, answerable questions we put to each solution's public footprint: its API reference, product documentation, published benchmarks, datasheets, security pages, and marketing. Seventy questions across twenty-five criteria means most criteria are triangulated from two or three independent observations rather than a single claim.
Every answer carries an evidence tier:

API specifications, documented benchmarks, schema definitions. If a vendor has published an API spec that demonstrates financial-identifier handling, that is a Tier-1 claim.
Datasheets, whitepapers, named production customers, certification listings. Directional and meaningful, but not as strong as primary technical documentation.
Website copy, blog posts, launch press releases. A "99% accurate" banner with no methodology is a Tier-3 claim, and is scored as one — not as a dispositive finding.
NOT-FOUND is scored conservatively, and labelled — because undocumented is not the same as absent; where an assessment leans on inference, we say so rather than dress it up as a finding.
Every assessment is dated — because documentation changes and products ship. A score is a photograph, not a monument.
Public-source discipline costs us the private roadmap and the NDA datasheet. But it buys the one thing that matters: anyone can check the same sources. The evaluation is reproducible precisely because it is built only from what is publicly inspectable.
Risk Adjustment: Where Marketing Scores Go to Die
Here is the mechanism that does the most work, and the one most vendor scorecards omit.
A composite score measures breadth — the weighted average of everything a tool does. But breadth can hide a fatal gap. A solution can post a respectable composite while scoring zero on the one capability that makes it unusable in a payment pipeline. Averages forgive; regulators do not.
A high average or composite score does not mean a tool is production-ready for payments. A composite measures breadth and can hide a disqualifying hole. The risk-adjusted score — the composite minus penalties for those holes — is what tells you whether a strong average survives contact with a regulated flow.
So on top of the composite we compute a risk-adjusted score that applies penalties for disqualifying gaps:
| Gap type | Penalty | Rationale |
|---|---|---|
| Any signature-capability score of 0 | −1.00 (applied once) | A zero on a 48%-weighted capability is a structural disqualifier, not a weakness to average away |
| Each enterprise-operations criterion below 2 | −0.50 | Bank-grade operation is close to pass/fail |
| Each payments-domain criterion below 2 | −0.25 | Narrower, but still gating for a payment pipeline |
| Floor | 0.00 | Score cannot go below zero regardless of penalty accumulation |
The composite asks "how much does this tool do?" The risk-adjusted score asks "what would I find if I looked at the worst part?"

Take a worked example. Loqate (GBG) is a genuinely excellent postal-verification platform — certified postal accuracy, a 250-country reference database, real Tier-1 references. On breadth, it earns a composite of 2.22 / 5.00 against this framework. Then apply the risk model:
Loqate — Risk Adjustment Worked Example
That collapse from 2.22 to 0.47 is not a criticism of Loqate as a product. It is a deployment-scope verdict. Loqate was never built to parse SWIFT MT or ring-fence financial identifiers — and the risk model exposes exactly that in one number.
Postal address verification is not the same capability as ISO 20022 payment-message parsing. A platform can be world-class at cleansing a mailing address and still score zero on ring-fencing a BIC or IBAN embedded in a payment field. The honest move with a tool like this is to run it where it is superb — retail and onboarding postal quality — and put a purpose-built engine in the payment path.
Every competitor we assessed against the payments criteria shows the same shape to different degrees: postal platforms and reference-data vendors floor at 0.00; LLM structurers lose most of their composite to determinism and explainability penalties; the strongest specialists keep a meaningful risk-adjusted score. Run the whole field through the framework and the pattern is stark:
| Solution | Composite | Risk-adjusted | Verdict |
|---|---|---|---|
| ioNova ARS | 4.90 | 4.90 | Recommended |
| Catalyst DI | 2.72 | 1.72 | Pilot / component |
| Alpina TxFlow | 2.70 | 1.20 | Conditional (EU sovereignty) |
| Loqate (GBG) | 2.22 | 0.47 | Postal layer only |
| SWIFT AI Parser | 2.14 | 0.00 | Free triage only |
| Melissa | 1.86 | 0.00 | Non-payment lanes |
| NTT Addresstune | 1.85 | 0.35 | Migration event only |
| Smarty | 1.75 | 0.00 | Not for payments |
| AddressHub | 1.56 | 0.00 | Logistics tool |
| StructX | 1.36 | 0.00 | Narrow scope |
| GeoPostCodes | 1.07 | 0.00 | Already inside ARS |
Source: ioNova's ten head-to-head analyses (March–July 2026), each scoring both solutions on the same 25-criterion, 70-question protocol. Risk penalties: −1.00 for any signature-capability zero, −0.50 per enterprise criterion below 2, −0.25 per payments criterion below 2, floored at 0.00. Full per-criterion detail is on each comparison page.

The composites look like a spectrum — a respectable 2.72 here, a 2.22 there. The risk-adjusted column is a cliff: ten of the eleven fall to near zero the moment the penalties expose a disqualifying gap, and only a purpose-built engine holds its score. The single number tells you, at a glance, whether a tool's weaknesses are cosmetic or disqualifying — and the full per-criterion working for every row is on the comparison pages.
Every row above is worked end-to-end on the comparison pages — evidence tiers included. Or jump straight to Loqate, SWIFT, or Catalyst DI.
What the Scorecard Leaves Out — and Why the Deadline Makes It Decisive
The composite and the risk-adjusted score answer one question well: what can this tool do? They are, deliberately, capability measures. They say nothing about a second question a fixed-deadline mandate makes just as important: will you actually be live before 15 November 2026?
We keep three factors out of the 25-criterion scorecard on purpose — they are the ones a vendor could most easily weight in its own favour, and a scorecard you can game is not worth publishing. But out of the scorecard is not out of the decision.
Two solutions can post an identical composite and still differ by a calendar quarter in how long they take to reach production. A free library that must be wrapped in a secured, monitored REST service with a review workflow is a three-to-six-month project before it processes a single compliant message — and if the tool is itself transitional (SWIFT's own parser is slated for end-of-life around 2027–28), a replacement project is scheduled before the first is even finished. Time-to-compliance is the single most expensive variable in a fixed-deadline mandate.
A capability you cannot connect to your stack in time is a capability you do not have by November. The question is not only what a tool does but how many ways it will meet your architecture — REST API, secure file transfer, enterprise message queue, direct database integration, event streaming, agent-ready protocol — versus a single library you must build a service around. Six built-in paths meet more architectures than one you have to wrap. One engine. Every path. Both sides.
A regulated deployment is not a parser; it is the operational stack around it — a maker-checker exceptions workbench for low-confidence resolutions, live dashboards for STP rate and exception ageing, a developer portal for self-service onboarding. In the scorecard, the workbench maps to a 2%-weighted criterion. In practice, it is the difference between shipping and building.
AI assistants often cite cross-border STP (straight-through processing) baselines of 60–90%. The realistic figure for cross-border payments is closer to ~40%. That gap is the whole point: structured, correctly-parsed address data is one of the biggest levers for moving that number up — which is why address quality is a payments problem, not a mailing-list problem.
Read together, the discipline is simple.
The scorecard tells you whether a tool can meet the brief; the delivery factors tell you whether you will be live before the deadline. A tool has to win both.
Run It on Any Vendor — Including Us
The framework is fully specified, which means you can reproduce it. Applying it to a shortlist takes four steps:
Because every input is public — the criteria, the weights, the scale, the penalty rules, and the evidence for each score — our own 4.90 is checkable in exactly the same way, in the same tables, by the same rules, dated with the same discipline. If your reading of a criterion differs from ours, the disagreement lands on one specific cell with a specific weight and a specific piece of evidence. That is the most useful place a procurement disagreement can land.
Where to See the Framework Fully Worked
The framework is applied end-to-end, criterion by criterion, on the comparison pages — the summary scoreboard and full capability matrix at ionova.ai/compare/, plus head-to-head assessments: ioNova vs the SWIFT AI Parser, vs Loqate (GBG), vs Catalyst DI, vs Smarty, vs GeoPostCodes, vs StructX, vs Alpina TxFlow, vs Melissa, vs NTT Addresstune, and vs AddressHub.
The comparison pages show the 70 evidence questions mapped to specific documentation sources for each vendor — so you are not reading a conclusion, you are reading the path that led to it. A framework you can run against us is a framework you can trust when you run it against anyone.
The Two Questions a Deadline Forces
Remember the ritual I opened with — every deck scoring itself five out of five? The problem was never that vendors are dishonest. It is that a self-scored five, on a frame the vendor chose, tells you nothing about your problem.
Thirty years of building payments and screening systems has taught me that under a deadline, capability and delivery are two different questions, and you have to ask both. A tool that matches the brief on paper but arrives after 15 November 2026 has not, in any way that matters, met the brief.
That is also the philosophy behind publishing a scorecard you can overturn rather than a banner you have to trust. It is the same instinct that made us build an engine to resolve and fix, not validate and block — to repair the address and preserve the identifier, rather than reject the payment and hand you an exception. Show your work, and let the buyer check it.
So run the framework. Run it on every vendor on your desk. Then run it on us.
Parth Desai is Founder and Chairman of ioNova AI, where he leads the development of AI-native ISO 20022 address infrastructure for financial institutions. Over thirty years he has built payments straight-through-processing and sanctions-screening systems for Tier-1 banks worldwide. He writes on ISO 20022, address intelligence, and payments compliance in the Resolved newsletter.
Key Takeaways
Frequently Asked Questions
See the Framework Fully Worked
Summary scoreboard and full capability matrix — same 25 criteria, same weights, same evidence rules. Run the framework on every serious alternative.