Network Codify

← Blog

Cardinality: what a dimension costs, depending on the architecture

When you build a network observability platform, you naturally want to add context to the data you collect. Take a metric or a syslog event about an interface: one could consider enriching it with the VLAN it carries, the subnet, the VRF, and even the name of the application concerned, depending on what the environment can associate with that interface. The cautious rule is well known: enrich at ingestion with stable, unambiguous dimensions only, and leave the rest to read time. This article explains why that caution is not a matter of principle, and why its price varies with the system that stores the data.

Not every useful piece of information should become an identity dimension. Where it sits in the data model matters as much as what it contains, and to understand why, you have to start from a simple notion: cardinality.

What cardinality is

The cardinality of a field is the number of distinct values it takes in a dataset. For sizing a platform, what matters is less that number at a given moment than the number it could reach. An environment field that is production, staging or development has a cardinality of three, even if a million logs carry the value production.

Information Example values Cardinality to expect
Environment production, staging, development Low and bounded
Vendor cisco, arista, juniper Low
Device router01, router02, switch01 Tied to the size of the estate
Source IP address 10.1.2.34, 192.0.2.18 Potentially very high
Session identifier 8af923, b710ef Grows with every new session

Cardinality depends on scope and period. An IP address that identifies a device in a stable estate does not have the same profile as a client address on an Internet-facing service. And high cardinality does not mean “a lot of data”: it means many different values, or many different combinations. To understand why these combinations matter, we need to look at which dimensions identify the stored data.

When a dimension participates in identity

A dimension is a name and value pair: device="router01", site="paris", environment="production". It is used to select and to group: the metrics of one router, the logs of one site. In systems that organise data into series or streams, some dimensions also play a structural role: they participate in the identity of the group in which data is filed. Let us call them identity dimensions. Other information can remain in event content, where the model allows it. Logs and metrics offer different possibilities: the examples below show what each model allows us to retain and search.

The simplest image is a filing cabinet. Adding a measurement or a log slips one more sheet into an existing folder. Introducing a new combination of identity dimensions opens one more folder. A cabinet with many sheets in a few folders is not managed like a cabinet with millions of nearly empty folders. Identity dimensions pick the folder; data fills it.

Two consequences follow from this image, before any product is named. First, what matters is not the cardinality of each dimension taken on its own, but the number of combinations actually observed, and the speed at which new combinations appear. A thousand devices with a hundred interfaces make a hundred thousand folders; adding a dimension with a thousand values per combination makes a hundred million, in theory. Only the combinations actually received exist. And a dimension whose value follows from another adds no combination at all: a thousand devices that each have a single site still make a thousand combinations when the site is added, not a thousand times the number of sites. This reasoning assumes that the association stays stable over the period studied.

Second, churn matters as much as count: a dimension that takes a new value with every event, a session identifier for instance, keeps opening folders that will only ever hold one sheet. Those are two different loads. A lot of data in stable folders costs volume: storage, predictable, that compresses well. A continuous creation of folders adds work per identity: index entries, memory, and sometimes sparsely filled data chunks. Managing and retaining them has a cost, even if few folders are open at the same time. How those chunks are grouped, compacted and deleted depends on the engine.

The remaining question is where, in the collection chain, the decision that opens folders is made.

The problem is not extracting a field, it is making it an identity dimension

A device sends a syslog message whose content looks like LOGIN_SUCCESS user=alice src_ip=10.1.2.34 session=8af923. A collection pipeline, whatever the tool, parses that text, normalises the field names and adds context. The event becomes a structured object:

{
  "device": "router01",
  "site": "paris",
  "severity": "info",
  "event_type": "LOGIN_SUCCESS",
  "username": "alice",
  "src_ip": "10.1.2.34",
  "session_id": "8af923"
}

Extracting username, src_ip or session_id is useful, even though those fields have many distinct values. And that parsing, on its own, opens no folder: it enriches the event. The decision that matters comes later, when the event is sent to storage: which pieces of information become identity dimensions, and which stay attached to each event. That choice is configured in the collector, in its output component or in the ingestion chain.

Parsing produces a structured event. Only the fields turned into identity dimensions open folders; the rest travels in the content.

Keeping session_id in the content of each event does not increase the number of folders. Making it an identity dimension brings its values into the identity of the groups, and every new session then opens a combination. That does not make the pipeline free: parsing consumes CPU, extra fields make events larger, stateful processing or per-session aggregation consumes memory. But those costs are of a different nature from the explosion of the number of folders. And this is where the architecture of the storage system comes in: not all of them have folders, and those that do pay for them at different prices.

What it costs, depending on the architecture

Metrics in series (Prometheus, InfluxDB 1 and 2)

Let us start with the systems to which the filing-cabinet image directly applies. Prometheus and its derivatives (Mimir, Thanos, VictoriaMetrics) store time series: each series is identified by a metric name and a set of labels, and holds timestamped samples.

interface_receive_bytes_total{device="router01", interface="Eth1"}
series identity: metric name + labels
    10:00 → 123456
    10:01 → 128000
    10:02 → 135000
the counter changes at every scrape, the series stays the same
As long as the identity does not change, the scrapes feed the same series.

Two devices with two interfaces each make four series: the number of series follows the size of the estate, nothing to worry about. The problem appears when an unbounded dimension enters the identity, a source address on a packet counter for instance.

packets_total{device="router01", interface="Eth1", src_ip="10.1.1.10"} 42
packets_total{device="router01", interface="Eth1", src_ip="10.1.1.11"} 17
every address seen on this interface opens a series
1,000 devices × 100 interfaces × 2 directions = 200,000 series
× 1,000 addresses per combination = 200,000,000 series
A scenario, not a forecast: only the combinations actually received become series.

Every series means work: maintaining its identity, handling its samples, finding it in the indexes. An excessive multiplication is paid in memory, in storage and indexes, in ingestion cost, in slower queries and alerting rules as soon as they scan many series, and in the extreme in saturation. Churn, the continuous creation of short-lived series, is a load of its own, and controlled cardinality does not make volume free: scrape frequency and retention still count. That is why Prometheus best practices advise against unbounded dimensions, user identifiers first among them.

InfluxDB, in versions 1 and 2, follows the same model under other names: a series is a measurement plus a set of tags, tags are indexed, fields are not. The tag versus field split is exactly the identity versus content split, and the explosion of the number of series when an IP address is turned into a tag is an accident many network teams have lived through.

Separating context from measured detail

Limiting metric labels raises a question: what information can be recovered later? If you only keep a packet counter per interface, no join will ever give back the breakdown of those packets by source address. That information was never measured; it is gone.

So two things the word “dimension” conflates have to be told apart. Descriptive context, a device’s site, its role, its VRF, can be recovered by joining with another source, the inventory for instance, and does not need to enter the identity of the series. A measurement dimension, the source address of a traffic flow, the queue of a counter, is part of the observation: removing it from the identity means giving up the detail, and that is decided knowing the analyses expected, not to save series.

Let us look at how to separate that descriptive context from the measurement in Prometheus. In its usual series model, a sample does not carry arbitrary context fields comparable to log content: that context is generally represented by labels, which participate in the identity. Joining context at read time takes two forms there. The first is the info series, what Prometheus calls an info metric, by convention suffixed _info with a constant value of 1: a context series with controlled cardinality, for instance interface_info{device, interface, vlan, vrf} 1, generated from the source of truth, and joined to the measurement at query time:

rate(interface_receive_bytes_total[5m])
  * on (device, interface) group_left (vlan, vrf)
    interface_info

The VLAN and the VRF do not enter the counter’s identity; they live in a separate series, and changing the VLAN or VRF creates a new info series. Churn is concentrated on those series without changing the counters’ identities. The join requires exactly one matching interface_info series per (device, interface) pair at each evaluation time. The source must stop exposing the old combination when it exposes the new one; multiple matches cause this join to fail. This example therefore assumes an unambiguous VLAN and VRF per interface.

The second form is the source of truth producing the selector: you first ask it which interfaces carry the application you are looking for, and its answer becomes the filter on device and interface of the PromQL query. The inventory filters, the TSDB measures.

In both cases, recovering past context requires retaining its history. A router that moves from Paris to Lyon keeps its old metrics; joining them with its current site attributes to Lyon traffic produced in Paris. You have to decide whether the analysis is about the context at the time of the event or the current context, and in the first case the context source must keep validity periods and allow a temporal join. The same goes for correlated dimensions: “a device has only one site” is true at a given moment, not necessarily over the whole retention, and adding the site then creates new identities with every move.

Logs in streams (Loki)

Loki applies the same principle to logs. A stream is identified by its set of labels, within one tenant, and holds timestamped entries. The text of the message can change completely from one entry to the next without creating a new stream; the labels define its identity. Take three events with the same device and severity labels but three different source addresses.

{device="router01", severity="warning"}
    ├── 10:00:01 LOGIN_FAILED src_ip=10.0.0.1
    ├── 10:00:04 LOGIN_FAILED src_ip=10.0.0.2
    └── 10:00:08 LOGIN_FAILED src_ip=10.0.0.3
src_ip in the content: one stream, three entries

{device=“router01”, severity=“warning”, src_ip=“10.0.0.1”} {device=“router01”, severity=“warning”, src_ip=“10.0.0.2”} {device=“router01”, severity=“warning”, src_ip=“10.0.0.3”} src_ip as a label: three streams, one entry each

The number of logs has not changed. Their distribution has.

Loki indexes the labels of its streams and stores the logs in compressed blocks, the chunks. When logs scatter across a multitude of very short streams, the index grows and the chunks multiply, which raises costs and degrades performance. Nor is the goal a single stream, since throughput per stream matters too: you want useful, stable groupings.

Searching fields kept outside the labels

Let us return to the session_id extracted from syslog. Kept in the content, it creates no additional stream, but it must still be useful in an investigation. In Loki, if the stored lines are JSON, a LogQL query selects streams by their labels, parses their lines and filters on that field:

{device="router01", severity="info"}
  | json
  | session_id="8af923"

The braces select streams from the indexed labels. The rest extracts the fields and keeps only the events of the session being looked for. This read-time extraction does not modify the stored streams; it has a processing cost, which you contain by bounding the scope and the period. To retain information outside stream identity, Loki also offers structured metadata, which attaches information to entries without including it in the log text or in the stream labels.

Place Example Opens a folder?
Identity dimension device="router01" Yes
Field in the content "session_id": "8af923" No
Structured metadata session_id attached to the entry No

Keeping a piece of information in the event guarantees that it exists, not that a search will find it fast enough to be useful during an incident, when the engineer is waiting for the answer on screen. The example above starts from a favourable case: the device and the severity are known, the streams to scan are few. In an investigation, you often know nothing but a session identifier, with no device, no site and an uncertain period, and the same query then has to scan a much larger share of the logs. The design question most often missing is this one: what information does the investigator start with, over what period, and how quickly do they expect an answer? That is what decides whether a field can stay in the content, or needs an index. Even structured metadata has a downside: counting or grouping logs by a high-cardinality metadata field recreates, for the duration of the query, as many groups as there are values, and Loki caps that number. The cardinality kept out of the identity then reappears at read time. Keeping the field therefore addresses its availability; the search method determines the work needed to use it.

Columnar stores (InfluxDB 3)

The cost of an index per combination is not shared by every engine. The move from InfluxDB 1 and 2 to InfluxDB 3 helps explain the difference.

InfluxDB 3 and the columnar warehouses used behind some stacks arrange the data differently: a dimension is a column, every row carries its values, and a combination of dimensions no longer has an index entry of its own. Series identity has not disappeared, InfluxDB 3 keeps a primary key made of the timestamp and the tags, but it no longer materialises one folder per combination. The series cardinality limit, which was a cliff in InfluxDB 1 and 2, therefore largely disappears in InfluxDB 3.

Concretely, you no longer name a series, you filter for it. Each measurement is a table, tags and fields are columns, and the series of one interface is read in SQL like any other rows:

SELECT time, rx_bytes
FROM interface
WHERE device = 'router01'
  AND interface = 'Eth1'
  AND time >= now() - INTERVAL '1 hour'
ORDER BY time;

Tags still serve to filter. Persisted data is stored in Parquet files, a columnar format. The engine can use sorting and statistics to skip data that a query does not need. How effective this selection is depends on data organisation and the filters used: a filter on src_ip alone does not guarantee that little data will be read.

Take a src_ip tag containing a million distinct addresses. In InfluxDB 1 and 2, those addresses can multiply the series and increase the size of their index. InfluxDB 3 removes this constraint associated with the older storage engine.

Cardinality still matters when considering query costs. Looking up measurements for a single address requires different work from calculating a total for each address: in the second case, the engine may need to calculate and return a million results. Even a targeted search can read a lot of data if its organisation does not allow the engine to skip irrelevant files effectively.

InfluxDB 3 therefore lets you retain high-cardinality dimensions, but query speed still depends on the data to scan and the results to produce.

Inverted indexes (Elasticsearch, OpenSearch and Splunk)

For logs, another approach prepares searches by indexing values present in events. That is the principle to examine with Elasticsearch and OpenSearch. An inverted index works like the index at the back of a book: for every value, the list of documents that contain it. Instead of starting from a document to read its fields, the engine starts from a value to find its documents, hence the name. By default, every field of a document enters that index; the mapping, that is the schema describing the fields, allows some to be excluded or indexed differently. Searching for 10.1.2.34 then means opening the 10.1.2.34 entry and reading the list of documents, without scanning the logs: finding the documents that contain one precise address is what these engines are built for. For logs, there is no group identity: it is the indexed fields and the aggregations that carry the cost.

That cost is paid in three places. First in aggregations: grouping documents by a field containing a million distinct values can require substantial memory and computation. The cost depends on the aggregation, the volume of data scanned and the number of groups to produce. Then in the size of the index: every indexed field adds its entries, and a log rich in thirty fields weighs far more than the text it contains. Finally in the structure itself: if a pipeline creates a new field per value, for instance a field named after the session identifier, the schema grows with every event, and that is the case that brings a cluster to its knees, long before a single field with a million values. The choice analogous to “identity or content” is therefore made field by field: index this field or not, and allow aggregations on it or not.

Splunk also indexes event text, but the choice of when to extract fields distinguishes it from Elasticsearch and OpenSearch. Beyond the fields indexed by default, it can extract fields at search time. An extraction can therefore evolve over retained events, at the cost of the processing needed to parse them. Index-time extraction remains possible, with additional indexing and storage costs. This trade-off therefore deserves its own row in the comparison.

Comparing advantages and limits

The preceding examples show the trade-offs between storage, ingestion and search. The table brings them together to compare the approaches. Metrics and logs are separated to make those trade-offs explicit. A single engine can combine several mechanisms; this table is not a performance ranking.

Approach Concrete advantages Drawbacks and limits
Metrics in series: Prometheus, InfluxDB 1 and 2 Compact storage for repeated measurements. With stable dimensions, many samples share the same identity and compress within the series. Every new combination adds index and memory costs. Highly variable labels create many short-lived series and can exhaust resources, even with few samples per series.
Logs in streams: Loki Less data to index than with extensive field indexing. With controlled labels and well-filled streams, the small index and compressed logs reduce storage costs. Searches outside the labels can be slow. Finding a session without knowing the device or site may require scanning many logs. Making every session a label shifts the problem: the index grows and small chunks multiply.
Columnar storage: InfluxDB 3 Retain detail without the old cardinality obstacle. Measurements per address or flow can be kept without multiplying series-index entries as in InfluxDB 1 and 2. High cardinality still costs work to analyse. A total per address can produce a million groups. A targeted search can also remain slow if data organisation requires reading many files.
Indexed fields: Elasticsearch, OpenSearch Find a value in an indexed field faster than by scanning all logs in scope. An address, user or session can be the starting point, even without knowing the device. That search gain is paid for at write time and in storage. Building indexes adds ingestion work, and retaining them takes space. They do not make aggregations over millions of values free.
Indexed text and search-time field extraction: Splunk Explore fields without defining them all at ingestion. An extraction can be added or changed and applied to retained events, provided their content contains the information. Extraction adds work to searches that use it. Across a broad scope, parsing events can increase response times. Indexing more fields adds ingestion work and increases index size, without guaranteeing faster searches.

Before choosing a label, check how many combinations it can create and how quickly they change, its usefulness for searches, and the detail you would lose by leaving it out.

The rule to keep

For logs, extracting a field costs nothing by itself. Where it is kept, in the content, in a field index or in the identity of a group, determines its effect on cardinality and search costs.

For metrics, adding a label changes series identity and can multiply their combinations.

What is left out of the identity must remain retrievable. A field left in log content must be searchable within a delay acceptable for the investigation. Context recovered through a join, such as a device’s site, must be the one that was true at the time of the event, not today’s.

The decision therefore has two levels: retain the information needed for analysis, then decide how to store and retrieve it. For logs, a useful field can remain in the content without becoming a label. For metrics, a measurement dimension that does not enter series identity at collection time is lost: no join will recreate it later. Only descriptive context can be supplied by another source, with the necessary history.


Recommended reading: the Prometheus data model and its naming best practices, the Loki documentation on cardinality and structured metadata, and the Elasticsearch documentation on mapping.

Additional references: InfluxDB 3 schema design, Prometheus storage and vector matching.

See also Loki’s indexing model, Elasticsearch aggregation structures and Splunk field extraction.