Time-Series Storage: How to Evaluate Encoding and Compression for IoT Data
DEV Community

Time-Series Storage: How to Evaluate Encoding and Compression for IoT Data

Encoding versus compression

Encoding chooses a representation for values before they are stored. Time-series data often has useful structure: timestamps follow a time axis, neighbouring values may be similar, and each measurement has a stable type and device context. Compression then reduces redundancy in that representation. A simplified pipeline looks like this:

timestamped measurements -> encoding -> compression -> stored time-series blocks -> decode/query when read

The goal is not merely to make a file smaller. The stored representation still has to support continuous writes, historical reads, aggregation, recovery, and whatever operational processes depend on the data.

Start with the data profile

Do not select a setting from a single compression-ratio example. First describe the series you actually expect to store:

  • devices: number and hierarchy
  • measurements: temperature, pressure, vibration, counters, status, ...
  • sampling frequencies: per measurement group
  • arrival pattern: steady, batched, bursty, or mixed
  • event-time disorder: expected late-arrival pattern
  • retention: recent and historical windows
  • queries: latest values, ranges, aggregations

Different signals can behave very differently. A slowly changing temperature series may contain more repeated structure than a noisy vibration series. A counter has a different pattern again. A realistic sample should include the main types rather than replacing them with random values.

Match the method to the signal

Apache IoTDB V2.0.x provides encoding methods for different data types and patterns. Three examples make the selection logic concrete:

  • TS_2DIFF is suited to monotonically increasing or decreasing integer sequences.
  • RLE is useful when values repeat consecutively.
  • GORILLA is designed for nearby consecutive numeric values and is a better fit than differential or run-length approaches for some floating-point series.

Compression follows encoding and operates on the binary stream. V2.0.x supports several compression methods, including LZ4, SNAPPY, GZIP, ZSTD, and LZMA2. Treat them as candidates to test rather than a fixed best-to-worst list.

What to measure

A useful experiment changes one major storage variable at a time and records the surrounding conditions. At minimum, capture:

  • stored size for the same input data
  • write acceptance and end-to-end availability
  • CPU and memory during writes and reads
  • recent-value and time-range query behavior
  • aggregation behavior over longer windows
  • compaction, recovery, backup, or retention effects relevant to the deployment

Run representative queries while ingestion continues. A setting that reduces disk use but makes routine historical queries impractical may not be a good fit. Likewise, a setting that looks good in a short load may behave differently after the system has accumulated a realistic history.

A small, reproducible test plan

The following process is intentionally release-neutral. Exact configuration names and supported options should be checked in the Apache IoTDB documentation for the version being evaluated.

  1. Prepare a fixed sample
    Export or generate a bounded sample with known device identifiers, measurement types, event timestamps, and values. Keep the sample unchanged across runs.

  2. Include realistic timing
    Mix the expected sampling frequencies. If the ingestion path can receive buffered or retried records, include a controlled amount of late data while preserving event time.

  3. Establish a baseline
    Load the sample with the baseline configuration. Record the configuration, software version, deployment shape, hardware limits, duration, and client behavior.

  4. Evaluate storage and reads together
    Measure the stored footprint, then run the same latest-value, range, and aggregation queries. Repeat selected reads while new data is being written.

  5. Compare one change at a time
    Change one encoding or compression-related option, rerun the same workload, and compare the complete result. Avoid changing schema, batch size, retention, and storage settings in the same run unless the experiment is explicitly testing that combined design.

How Apache IoTDB fits

Apache IoTDB is an Apache open-source, IoT-native time-series database. In its Tree Model, each time series can have its own data type, encoding method, and compression method. The project supports multiple encoding methods for different data types and applies compression after encoding, providing a storage and query foundation in which those choices can be evaluated alongside ingestion and time-aware analysis.

That relationship is important for developers. Storage is not a post-processing step detached from the rest of the system. Device identity, measurement meaning, event time, late arrivals, and query windows all influence whether a storage design remains useful. Before using a particular option, verify its exact name, default, supported data types, and version behavior in the release documentation. Then test with the workload you intend to operate.

Conclusion

Encoding and compression are valuable because time-series data contains structure that generic storage may not exploit efficiently. But storage efficiency is only one part of the result. The best configuration is the one that balances footprint, ingestion, resource use, query behavior, and operational work for a representative workload. To learn more and try Apache IoTDB, visit the Apache IoTDB GitHub repository.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.