Skip to main content

Optimizing system configuration

The number of possible setups is virtually unlimited, so this document cannot cover them all. Even similar builds and configurations can behave differently because of external factors. Your results may vary. Use these general guidelines as a starting point. Be cautious and make incremental changes. Test and observe each change before you move forward. Always focus on only one specific area at a time. Do not change the memory, storage, and CPU configurations all at once. Otherwise, it becomes nearly impossible to diagnose problems.

Memory management

These settings in /etc/sysctl.conf can optimize memory usage and disk I/O patterns:
To apply the changes, run sudo sysctl -p.

Network stack

These settings in /etc/sysctl.conf may improve network performance:

Storage configuration

For NVMe drives, optimize I/O scheduling: Storage optimization commands

Infrastructure monitoring

Monitoring is one of the most critical components of network infrastructure. This page covers monitoring, performance tuning, and alerting configuration for Cosmos SDK and Tendermint nodes.

Prometheus setup

First, install Prometheus:
This is an example Prometheus configuration:

Grafana integration

Install and configure Grafana:

Alert management

Install Alertmanager:

Log management

Loki setup

Use Loki for log aggregation:

Log rotation

Configure logrotate to manage log files:

Security configuration

Network security

Configure the UFW firewall:

Rate limiting

Validator-specific monitoring

Status query

Query the validator status through the SDK:
Query the status through the REST API:

Critical metrics

Monitor these validator-specific metrics:

Backup management

Host system monitoring

Resource usage tracking

Install and configure node_exporter:
Add node_exporter to the Prometheus configuration:

Seed node Prometheus metrics

Seed nodes now expose a Prometheus /metrics endpoint, so they can be scraped like any other node. Because a seed serves no RPC, this endpoint is the only observability surface it has — peer reachability and peer headroom both come from its p2p metrics (for example tendermint_p2p_peers). To enable it, set the usual instrumentation fields in config.toml under [instrumentation]:
Every exposed series is stamped with the node’s chain_id label, matching the behavior of full nodes.
For a seed node, a failure to bind the Prometheus listener (for example, the port is already in use or the address is unparseable) is fatal: the seed fails to start. This is deliberate — a seed’s value is only realized when it is observable, so an unambiguous crash is preferable to a seed that discovers peers while blind.A full node behaves differently: a bind failure is non-fatal. The full node logs the error and continues running, because it still serves RPC as an in-band way to observe its health.
The tendermint_p2p_peers gauge is present from node startup rather than appearing only after the first refresh interval, so an alert comparing peer count against the connection cap matches immediately rather than sitting unmatched during the startup window.

Metrics server timeouts

The Prometheus metrics HTTP server enforces request timeouts to bound how long a slow reader can hold a request slot:

max-open-connections behavior

The max-open-connections field under [instrumentation] is not a socket limit despite the name. It becomes promhttp’s MaxRequestsInFlight: the maximum number of scrapes served concurrently. Requests past the limit receive an HTTP 503 rather than being queued, and a slow reader occupies a slot until it finishes or hits the WriteTimeout. A value of 0 means unlimited. Keep this value above the number of scrapers that can overlap — for example, an HA Prometheus pair, blackbox probes, and a human running curl — so a burst of concurrent scrapes does not cause a healthy node to report as unscrapeable.

Cosmos SDK / Tendermint telemetry chain_id label

Node telemetry automatically injects a chain_id global label into the Cosmos SDK / Tendermint Prometheus telemetry (the GlobalLabels of the [telemetry] configuration) based on the node’s client chain ID. The label is appended only when a chain_id label is not already configured in GlobalLabels, so an explicit operator-set value always takes precedence. This is distinct from the chain_id constant label applied to the OpenTelemetry sei_chain namespace described above — it covers the separate Cosmos SDK/Tendermint telemetry path. Because every emitted series carries this label, update any PromQL that aggregates across series (for example sum(...) without (...) or by (...) clauses) to account for chain_id, and add a chain_id match where you want to scope a query to a single chain — useful when one Prometheus instance scrapes nodes on more than one chain.

New Tendermint internal metrics subsystems

Additional Prometheus metrics are exported under the tendermint namespace in these new subsystems:

EVM RPC OpenTelemetry metrics

The EVM RPC layer emits OpenTelemetry metrics through the process-wide MeterProvider (for example, a Prometheus exporter). It emits these metrics in parallel with the legacy sei_* metrics, so you can migrate dashboards incrementally.
Every OpenTelemetry metric series exported through the Prometheus exporter (the sei_chain namespace, including the evmrpc_*, flatkv_*, litt_*, and pebble_* metrics documented below) now carries a chain_id constant label derived from the node’s chain ID. The label is applied to every emitted series, so dashboards and alerting queries can group or filter by chain — for example when a single Prometheus instance scrapes nodes on more than one chain. Update any PromQL that aggregates across series (for example sum(...) without (...) or by (...) clauses) to account for the new label, and add a chain_id match where you want to scope a query to a single chain.

Available EVM RPC metrics

The evmrpc_request_latency_seconds histogram carries these labels:

Migrating from legacy EVM RPC metrics

These legacy sei_* metrics remain available today but are deprecated. They are scheduled for removal after dashboards migrate to the evmrpc_* OpenTelemetry metrics: Before the legacy metrics are removed, update your Prometheus and Grafana dashboards to use the evmrpc_* metrics. The latency unit changed from milliseconds (sei_rpc_request_latency_ms) to seconds (evmrpc_request_latency_seconds). Adjust any thresholds and panel formatting to match.

FlatKV OpenTelemetry metrics

The FlatKV state store emits OpenTelemetry metrics through the process-wide MeterProvider (for example, a Prometheus exporter). With these metrics, you can observe commit throughput, catchup progress, snapshotting, rollbacks, and snapshot imports.

Available FlatKV metrics

Labels

FlatKV metrics carry these labels where applicable:

Enabling Pebble internal metrics

A single FlatKV-level knob controls Pebble’s internal (per-DB) metrics: enable-pebble-metrics under [state-commit.flatkv] in app.toml (default true). The node honors this key when it is present, but the app.toml that seid init generates does not include it. To change the default, add the key manually. The value propagates to every data DB (account, code, storage, legacy, and metadata) during initialization. It overrides any per-DB EnableMetrics settings. Configure Pebble metrics through this knob, not through the individual per-DB settings.

Pebble estimated read/write metrics

SeiDB can emit lightweight, opt-in estimates of logical read and write operations observed by its Pebble-backed wrappers. These counters are disabled by default and are intended to approximate read amplification (estimated reads divided by estimated writes) without the overhead of Pebble’s full internal metrics.

Available metrics

Both counters carry a db attribute identifying the database the measurement applies to (derived from the base name of the data directory), so you can distinguish FlatKV data DBs, the state-store backend, and receipt storage.

Enabling the counters

These estimates are controlled per subsystem in app.toml. All default to false: For example, to enable the FlatKV counters, add the following to app.toml:

Tuning littidx eth_getLogs parallelism

When the receipt store uses the littidx backend, each eth_getLogs query scans the requested block range across a bounded worker pool instead of one block at a time. Per-block tag index scans and litt body reads are independent, so fanning them across multiple workers reduces latency on wide-range queries while results stay in strict (block, txIndex) order. The log-filter-parallelism config key under [receipt-store] in app.toml bounds how many blocks a single query scans concurrently. It defaults to 16 and applies only to the littidx backend. A value <= 0 falls back to the default.

Per-block hash logger

The state-commit store can record a per-block CSV of named block hashes to disk as a debugging and forensics tool. For each committed block it logs the memIAVL per-module and root hashes, the flatKV per-DB and root hashes, the application hash, the Tendermint block hash, the result hash (the merkle root over the block’s deterministic transaction results, equal to the next block’s LastResultsHash), and the changeset hash. Recording the same hashes under one block number on every node makes it straightforward to pinpoint where and when block-hash computation diverges across nodes. The feature is enabled by default. By default it writes files into a hash.log directory under the state-commit store’s data directory (that is, <home>/data/hash.log).

Configuration

The hash logger is configured with five app.toml fields under [state-commit]: Retention is disk-driven by default: up to roughly 16 GiB of sealed files are kept, with block-count retention disabled. If both sc-hash-logger-blocks-to-retain and sc-hash-logger-max-disk-size are set to 0, no retention bound applies and the logs grow without limit — a deliberate operator choice. For example, to disable the hash logger, add the following to app.toml:

LittDB OpenTelemetry metrics

LittDB emits its metrics through the process-wide OpenTelemetry MeterProvider instead of a private Prometheus client. When MetricsEnabled is set, LittDB records its metrics into the global provider. How those metrics are exported depends on MetricsServeEndpoint (default false):
  • When MetricsServeEndpoint is false (the default), LittDB records into the already-configured global MeterProvider and leaves exporting to the embedding application. The application is responsible for standing up the Prometheus exporter and serving the registry. No exporter or /metrics server is created by LittDB, and MetricsPort is ignored.
  • When MetricsServeEndpoint is true, LittDB stands up its own Prometheus exporter on the global provider and serves /metrics on MetricsPort (default 9101).
MetricsPort is ignored unless both MetricsEnabled and MetricsServeEndpoint are true. The MetricsNamespace and MetricsRegistry config fields were removed, and all metric names use a fixed litt_ prefix.

Available LittDB metrics

| litt_compression_latency_seconds | Histogram | seconds | Latency of compressing a batch of values before they are written. Emitted only for tables with value compression enabled. | | litt_compression_uncompressed_bytes | Counter | bytes | The number of uncompressed value bytes submitted to compression since startup. | | litt_compression_compressed_bytes | Counter | bytes | The number of compressed value bytes produced by compression since startup. Compared against litt_compression_uncompressed_bytes, this gives the aggregate compression ratio and total bytes saved. | | litt_compression_ratio | Histogram | ratio | The per-batch compression ratio (compressed bytes divided by uncompressed bytes); lower is better. |

Attributes

Migrating from legacy LittDB metrics

The OpenTelemetry migration changed the metric names, units, and shape. You must update your existing Prometheus and Grafana dashboards:
  • Latency metrics moved from millisecond summaries (for example, {namespace}_read_latency_ms) to second histograms (litt_read_latency_seconds). Change thresholds and panel formatting from milliseconds to seconds.
  • Counters and gauges gained a fixed litt_ prefix and explicit units. For example, bytes_read became litt_bytes_read, and the cache weight gauge became litt_chunk_cache_weight_bytes.
  • The per-cache series that previously had separate metric names (for example, chunk_read_cache_* and chunk_write_cache_*) are now the shared litt_chunk_cache_* metrics. The cache attribute distinguishes them.
  • The MetricsNamespace and MetricsRegistry config fields no longer exist. Metric names are fixed, and the global OTel provider always backs the metrics. Set the scrape port with MetricsPort.

Performance testing

For specific customizations or other metrics, ask the Sei technical communities on Telegram or Discord.