
Start using accurate IP data for cybersecurity, compliance, and personalization—no limits, no cost.
Sign up for freeIn August we told database customers their files had shrunk by up to 80%, and promised to explain how. This is that explanation. The tool that does it is now open source.
We found a serious size saving in the MMDB file format, described below. We shipped it as an update to mmdbctl and as a standalone tool, mmdbshrink. The tool takes an existing MMDB file and writes a smaller one. The new file is a standard MMDB: every reader library that opens the old file opens the new one and returns the same answer for every IP address. Across the 89 databases we build every day, the catalogue went from 91.2 GB to 62.8 GB, a 31% reduction. The best single result was ipinfo_privacy, from 789 MB to 155 MB, an 80% reduction.
Nothing was compressed. There is no decompression step, no new reader, no new format. The saving is real at rest, on the wire, and in memory, because MMDB files are memory-mapped and the file size is the memory footprint.
This tool is available in https://github.com/ipinfo/mmdbctl/releases/latest.
An MMDB file has two parts. The search tree is a binary trie over the bits of an IP address. Start at the root, read the address one bit at a time, go left on 0 and right on 1, and stop when you reach a leaf. The leaf holds a pointer into the second part, the data section, where the records live: country, city, ASN, privacy flags, whatever the database carries.
The format's designers took care over the data section. It has a pointer type, so a record that repeats can be written once and referenced from many leaves, and the writers in common use do this.

That is why, in the ipinfo_lite run in the worked example further down, the data section does not change size at all.
The search tree got less attention. Every node is written out as two records, one per child, and a standard writer emits every node it visits. It never asks whether it has already written an identical node.
Consider two /24 networks in different parts of the address space that carry the same answer for every one of their 256 addresses. Below the /24 boundary, the subtrees under those two prefixes are identical, bit for bit: the same shape, the same leaves, the same pointers into the data section. A standard writer stores both.
Now scale that up. A privacy database is a small set of verdicts applied to millions of ranges. A location database has a few hundred thousand distinct city records spread over millions of prefixes. Wherever two prefixes resolve to the same record, and the same is true of everything beneath them, their subtrees are duplicates. In our files, a large share of the tree was duplicate subtrees.
Hash-consing is a long-established technique from compiler and symbolic-computation work, described by Filliâtre and Conchon (2006), Type-Safe Modular Hash-Consing, and closely related to the subgraph sharing in Bryant’s binary decision diagrams (1986). The idea is simple: keep a hash table of nodes already created, and reuse an existing node whenever its contents are structurally identical to the one you are about to create. Applied bottom-up to a trie, this means processing the children first, then sharing nodes with the same pair of children. The result is a directed acyclic graph: multiple parents can point to the same subtree, which is stored only once.
Applied to an MMDB search tree, the procedure is:
node_count and the encoded data references.This requires no change to the MMDB format. Each node contains two records, which can identify another node, a data record, or an empty result. Multiple parents can refer to the same node; the MMDB format specification already describes this kind of sharing for IPv4 aliases.
Sharing identical subtrees preserves the sequence of branch decisions for each IP address and leads to the same data record or empty result. The file’s layout and size change; its lookup results stay the same.

We were the first to identify and ship this saving. When we started, the reference writer emitted a fresh node per position, and so did the other writers we looked at. A conceptually similar technique appeared last year in a WhatsApp vulnerability paper, where the researchers needed to store a very large set of phone numbers and built their own custom structure to do it. Ours differs in one respect: the output is still a valid MMDB.
The reference writer emits a fresh node per position, and so do the other writers we looked at. A conceptually similar technique appeared last year in a WhatsApp vulnerability paper, where the researchers needed to store a very large set of phone numbers and built their own custom structure to do it. Ours differs in one respect: the output is still a valid MMDB.
Across our production catalogue:
Lookup speed is unchanged or slightly better. In the ipinfo_lite benchmark below, one million lookups ran at 797 ns per operation on the original and 757 ns on the optimized file, and resident memory fell from 36.02 MiB to 23.92 MiB. Fewer tree bytes means fewer page faults.
This is the tool run against our free Lite database, exactly as it appears in the README.
$ mmdbshrink --full-report ipinfo_lite.mmdb
input: ipinfo_lite.mmdb
output: ipinfo_lite.shrunk.mmdb
size: 37.76 MB -> 25.08 MB (saved 12.68 MB, -33.58%)
nodes: 3,790,310 -> 2,205,358 (-41.82%)
tree: 30.32 MB -> 17.64 MB
data: 7.44 MB -> 7.44 MB (unchanged)
elapsed: 452ms
validation (seed 1):
[OK] metadata ip_version=6 record_size=32 binary_format=2.0
database_type: ipinfo bundle_location_lite.mmdb build_epoch=1790582633
node_count: 3,790,310 -> 2,205,358 (-41.82%)
[OK] fixed probes n=39 hits=7 (0s)
[OK] random IPv4 n=1,000,000 hits=861,887 (86.19%) mismatches=0 (1.73s)
[OK] random IPv6 n=1,000,000 hits=1,985 (0.20%) mismatches=0 (39ms)
bench (1,000,000 lookups each):
input: size=37.76 MB per_op= 797ns qps=1,320,636
output: size=25.08 MB per_op= 757ns qps=1,319,347
memory (200,000 lookups each, separate processes):
input: mmap_cached=36.02 MiB proc_rss=29.34 MiB per_op= 707ns minflt=155 majflt=0
output: mmap_cached=23.92 MiB proc_rss=27.58 MiB per_op= 695ns minflt=168 majflt=0Read the last three lines together. The data section is untouched. Every byte saved came out of the search tree, and 41% of the nodes turned out to be duplicates of nodes already written.
A smaller file is worth nothing if a single lookup changes. Before any optimized file reached a customer we checked it four ways.
The first optimized files went into production in July, first on a customer's custom dataset that had grown past a 2 GB platform limit, then into our own API, where the smaller files cut the memory each instance needs. Today they are the default: 100% of API traffic and 100% of IPinfo database downloads serve optimized files under the same filenames as before.
MMDB is an open format that gives fast lookups of IP metadata, and the ecosystem supports it widely. We have published open source tooling for it before, and this continues that commitment. The duplication we found is a property of the format's writers, and anyone shipping MMDB files has the same tree bloat we had. What we compete on is what goes into the file. How the file is packed is not where our advantage lives, so the optimizer is open source under Apache 2.0.
If you build MMDB files, run it on yours. If you read MMDB files, you need do nothing at all.
You can find the project in this GitHub repository: https://github.com/ipinfo/mmdbctl/tree/master/mmdbshrink
You can easily install directly with go install github.com/ipinfo/mmdbctl/mmdbzip@latest, for other options see the Installation section in the README.md.
After installing you’re ready to use it like this: mmdbshrink ipinfo_lite.mmdb
If you are an IPinfo database customer, you don't need to do anything. Every MMDB you download already has these savings applied.

Tiago is Head of Anonymizer Detection, where he fine-tunes IP data streaming processes and transforms vast data sets into actionable insights. He was previously a staff research scientist at BitSight Technologies.

Silvano Cerza is an Integration Engineer at IPinfo, where he builds the bridges between IPinfo data and the platforms customers use, from Splunk and Snowflake to open-source SDKs and the MCP server for AI agents.