Skip to main content
2 hours ago by Tiago Martins & Silvano Cerza — 7 min read

How We Made Our MMDB Files a Third Smaller Without Changing a Byte of Data

How We Made Our MMDB Files a Third Smaller Without Changing a Byte of Data

Get Unlimited Access to IPinfo Lite

Start using accurate IP data for cybersecurity, compliance, and personalization—no limits, no cost.

Sign up for free

In August we told database customers their files had shrunk by up to 80%, and promised to explain how. This is that explanation. The tool that does it is now open source.

What We Did

We found a serious size saving in the MMDB file format, described below. We shipped it as an update to mmdbctl and as a standalone tool, mmdbshrink. The tool takes an existing MMDB file and writes a smaller one. The new file is a standard MMDB: every reader library that opens the old file opens the new one and returns the same answer for every IP address. Across the 89 databases we build every day, the catalogue went from 91.2 GB to 62.8 GB, a 31% reduction. The best single result was ipinfo_privacy, from 789 MB to 155 MB, an 80% reduction.

Nothing was compressed. There is no decompression step, no new reader, no new format. The saving is real at rest, on the wire, and in memory, because MMDB files are memory-mapped and the file size is the memory footprint.

This tool is available in https://github.com/ipinfo/mmdbctl/releases/latest.

How an MMDB File Is Laid Out

An MMDB file has two parts. The search tree is a binary trie over the bits of an IP address. Start at the root, read the address one bit at a time, go left on 0 and right on 1, and stop when you reach a leaf. The leaf holds a pointer into the second part, the data section, where the records live: country, city, ASN, privacy flags, whatever the database carries.

The format's designers took care over the data section. It has a pointer type, so a record that repeats can be written once and referenced from many leaves, and the writers in common use do this.

That is why, in the ipinfo_lite run in the worked example further down, the data section does not change size at all.

The search tree got less attention. Every node is written out as two records, one per child, and a standard writer emits every node it visits. It never asks whether it has already written an identical node.

Why That Matters: The Internet Is Repetitive

Consider two /24 networks in different parts of the address space that carry the same answer for every one of their 256 addresses. Below the /24 boundary, the subtrees under those two prefixes are identical, bit for bit: the same shape, the same leaves, the same pointers into the data section. A standard writer stores both.

Now scale that up. A privacy database is a small set of verdicts applied to millions of ranges. A location database has a few hundred thousand distinct city records spread over millions of prefixes. Wherever two prefixes resolve to the same record, and the same is true of everything beneath them, their subtrees are duplicates. In our files, a large share of the tree was duplicate subtrees.

The Fix: Hash-Consing the Tree

Hash-consing is a long-established technique from compiler and symbolic-computation work, described by Filliâtre and Conchon (2006), Type-Safe Modular Hash-Consing, and closely related to the subgraph sharing in Bryant’s binary decision diagrams (1986). The idea is simple: keep a hash table of nodes already created, and reuse an existing node whenever its contents are structurally identical to the one you are about to create. Applied bottom-up to a trie, this means processing the children first, then sharing nodes with the same pair of children. The result is a directed acyclic graph: multiple parents can point to the same subtree, which is stored only once.

Applied to an MMDB search tree, the procedure is:

  1. Traverse the existing tree from the leaves upward, processing children before their parents.
  2. For each node, replace references to child nodes with their canonical identities. References to data records and empty results retain their meaning.
  3. Use the ordered pair of child references as the key in a hash table. If that exact pair is already present, reuse the canonical node. Otherwise, add a new canonical node. The table checks the complete pair for equality, so hash collisions never cause a merge.
  4. Write out the canonical nodes, keeping the root at node zero, renumbering node references, and updating node_count and the encoded data references.

This requires no change to the MMDB format. Each node contains two records, which can identify another node, a data record, or an empty result. Multiple parents can refer to the same node; the MMDB format specification already describes this kind of sharing for IPv4 aliases.

Sharing identical subtrees preserves the sequence of branch decisions for each IP address and leads to the same data record or empty result. The file’s layout and size change; its lookup results stay the same.

We were the first to identify and ship this saving. When we started, the reference writer emitted a fresh node per position, and so did the other writers we looked at. A conceptually similar technique appeared last year in a WhatsApp vulnerability paper, where the researchers needed to store a very large set of phone numbers and built their own custom structure to do it. Ours differs in one respect: the output is still a valid MMDB.

The reference writer emits a fresh node per position, and so do the other writers we looked at. A conceptually similar technique appeared last year in a WhatsApp vulnerability paper, where the researchers needed to store a very large set of phone numbers and built their own custom structure to do it. Ours differs in one respect: the output is still a valid MMDB.

The Numbers

Across our production catalogue:

Database

Before

After

Reduction

ipinfo_privacy

789 MB

155 MB

80%

ipinfo_location

687 MB

316 MB

54%

ipinfo_core

1.5 GB

728 MB

53%

ipinfo_lite

38 MB

25 MB

34%

ipinfo_plus

5.6 GB

3.8 GB

32%

resproxy

3.5 GB

2.6 GB

27%

Total: 89 databases

91.2 GB

62.8 GB

31%

Lookup speed is unchanged or slightly better. In the ipinfo_lite benchmark below, one million lookups ran at 797 ns per operation on the original and 757 ns on the optimized file, and resident memory fell from 36.02 MiB to 23.92 MiB. Fewer tree bytes means fewer page faults.

A Worked Example

This is the tool run against our free Lite database, exactly as it appears in the README.

$ mmdbshrink --full-report ipinfo_lite.mmdb

input:   ipinfo_lite.mmdb
output:  ipinfo_lite.shrunk.mmdb

  size:   37.76 MB   ->  25.08 MB   (saved 12.68 MB, -33.58%)
  nodes:  3,790,310  ->  2,205,358  (-41.82%)
  tree:   30.32 MB   ->  17.64 MB
  data:   7.44 MB    ->  7.44 MB    (unchanged)

elapsed: 452ms

validation (seed 1):
  [OK] metadata        ip_version=6  record_size=32  binary_format=2.0
                       database_type: ipinfo bundle_location_lite.mmdb  build_epoch=1790582633
                       node_count: 3,790,310 -> 2,205,358 (-41.82%)
  [OK] fixed probes    n=39  hits=7  (0s)
  [OK] random IPv4     n=1,000,000  hits=861,887 (86.19%)  mismatches=0  (1.73s)
  [OK] random IPv6     n=1,000,000  hits=1,985 (0.20%)  mismatches=0  (39ms)

bench (1,000,000 lookups each):
  input:    size=37.76 MB    per_op= 797ns  qps=1,320,636
  output:   size=25.08 MB    per_op= 757ns  qps=1,319,347

memory (200,000 lookups each, separate processes):
  input:    mmap_cached=36.02 MiB   proc_rss=29.34 MiB   per_op= 707ns  minflt=155  majflt=0
  output:   mmap_cached=23.92 MiB   proc_rss=27.58 MiB   per_op= 695ns  minflt=168  majflt=0

Read the last three lines together. The data section is untouched. Every byte saved came out of the search tree, and 41% of the nodes turned out to be duplicates of nodes already written.

Proving the Output Is the Same Database

A smaller file is worth nothing if a single lookup changes. Before any optimized file reached a customer we checked it four ways.

  1. Full enumeration in both directions. Every prefix in the original is looked up in the optimized file and vice versa. For ipinfo_lite that is 3,783,922 prefixes, 7,567,844 lookups, all matching.
  2. Random probes. One million random IPv4 and one million random IPv6 addresses, zero mismatches.
  3. Existing tooling. mmdbctl verify reports the file valid, mmdbctl diff reports 0 subnets and 0 records modified, and MaxMind's own mmdbverify passes. We tested the files against 12 MMDB reader SDKs.
  4. Benchmark parity on latency and memory, as above.

The first optimized files went into production in July, first on a customer's custom dataset that had grown past a 2 GB platform limit, then into our own API, where the smaller files cut the memory each instance needs. Today they are the default: 100% of API traffic and 100% of IPinfo database downloads serve optimized files under the same filenames as before.

Why We Are Giving It Away

MMDB is an open format that gives fast lookups of IP metadata, and the ecosystem supports it widely. We have published open source tooling for it before, and this continues that commitment. The duplication we found is a property of the format's writers, and anyone shipping MMDB files has the same tree bloat we had. What we compete on is what goes into the file. How the file is packed is not where our advantage lives, so the optimizer is open source under Apache 2.0.

If you build MMDB files, run it on yours. If you read MMDB files, you need do nothing at all.

You can find the project in this GitHub repository: https://github.com/ipinfo/mmdbctl/tree/master/mmdbshrink

You can easily install directly with go install github.com/ipinfo/mmdbctl/mmdbzip@latest, for other options see the Installation section in the README.md.

After installing you’re ready to use it like this: mmdbshrink ipinfo_lite.mmdb

If you are an IPinfo database customer, you don't need to do anything. Every MMDB you download already has these savings applied.

Share this article

About the authors

Tiago Martins

Tiago Martins

Tiago is Head of Anonymizer Detection, where he fine-tunes IP data streaming processes and transforms vast data sets into actionable insights. He was previously a staff research scientist at BitSight Technologies.

Silvano Cerza

Silvano Cerza

Silvano Cerza is an Integration Engineer at IPinfo, where he builds the bridges between IPinfo data and the platforms customers use, from Splunk and Snowflake to open-source SDKs and the MCP server for AI agents.