Influx crashes regularly

Hi,

since yesterday, I’m observing regularly that influxdb crashes during compaction. I then always need to identify the malformed file and remove (or move it to a backup location) it.

The weird thing is that while there has definitely been no update, the issue appears again and again now (while I never experienced such issues in the past).

Am I the only one experiencing this?

Here is an except from the logs:

ts=2026-08-14T11:23:22.912759Z lvl=info msg="Deleted shard" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_delete_shard db_instance=c3efccc8b8ae0d98 db_shard_id=27675 db_rp=autogen
ts=2026-08-14T11:23:22.912832Z lvl=info msg="Deleting shard from shard group deleted based on retention policy (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_delete_shard db_instance=c3efccc8b8ae0d98 db_shard_id=27675 db_rp=autogen op_event=end op_elapsed=3.030ms
ts=2026-08-14T11:23:22.912930Z lvl=info msg="Dropping shard meta references (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_drop_refs db_instance=c3efccc8b8ae0d98 db_shard_id=27675 db_rp=autogen owners= op_event=start
ts=2026-08-14T11:23:22.923822Z lvl=info msg="Dropping shard meta references (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_drop_refs db_instance=c3efccc8b8ae0d98 db_shard_id=27675 db_rp=autogen owners= op_event=end op_elapsed=10.940ms
ts=2026-08-14T11:23:22.923947Z lvl=info msg="Drop phantom shard references (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_drop_phantom_refs db_instance=d64e712185312609 db_shard_id=27657 db_rp=autogen owners= op_event=start
ts=2026-08-14T11:23:22.923961Z lvl=warn msg="Expired phantom shard detected during retention check, removing from metadata" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_drop_phantom_refs db_instance=d64e712185312609 db_shard_id=27657 db_rp=autogen owners=
ts=2026-08-14T11:23:22.936591Z lvl=info msg="Drop phantom shard references (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_drop_phantom_refs db_instance=d64e712185312609 db_shard_id=27657 db_rp=autogen owners= op_event=end op_elapsed=12.661ms
ts=2026-08-14T11:23:22.936710Z lvl=info msg="Pruning shard groups after retention check (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_prune_shard_groups op_event=start
ts=2026-08-14T11:23:22.948152Z lvl=info msg="Pruning shard groups after retention check (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_prune_shard_groups op_event=end op_elapsed=11.445ms
ts=2026-08-14T11:23:22.948233Z lvl=info msg="Retention policy deletion check (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_event=end op_elapsed=411.041ms
ts=2026-08-14T11:53:22.538020Z lvl=info msg="Retention policy deletion check (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_event=start
ts=2026-08-14T11:53:22.538549Z lvl=info msg="Pruning shard groups after retention check (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_prune_shard_groups op_event=start
ts=2026-08-14T11:53:22.541166Z lvl=info msg="Pruning shard groups after retention check (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_prune_shard_groups op_event=end op_elapsed=2.621ms
ts=2026-08-14T11:53:22.541323Z lvl=info msg="Retention policy deletion check (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_event=end op_elapsed=3.325ms
ts=2026-08-14T12:00:40.896915Z lvl=warn msg="internal error not returned to client" log_id=14gzJF80000 handler=error_logger error="context canceled"
ts=2026-08-14T12:23:22.537452Z lvl=info msg="Retention policy deletion check (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_event=start
ts=2026-08-14T12:23:22.538053Z lvl=info msg="Pruning shard groups after retention check (start)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_prune_shard_groups op_event=start
ts=2026-08-14T12:23:22.540361Z lvl=info msg="Pruning shard groups after retention check (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_name=retention_prune_shard_groups op_event=end op_elapsed=2.321ms
ts=2026-08-14T12:23:22.540502Z lvl=info msg="Retention policy deletion check (end)" log_id=14gzJF80000 service=retention op_name=retention_delete_check op_event=end op_elapsed=3.082ms
ts=2026-08-14T12:46:58.503112Z lvl=info msg="Cache snapshot (start)" log_id=14gzJF80000 service=storage-engine engine=tsm1 op_name=tsm1_cache_snapshot op_event=start
ts=2026-08-14T12:47:07.854729Z lvl=info msg="Snapshot for path written" log_id=14gzJF80000 service=storage-engine engine=tsm1 op_name=tsm1_cache_snapshot path=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676 duration=9351.594ms
ts=2026-08-14T12:47:07.855896Z lvl=info msg="Cache snapshot (end)" log_id=14gzJF80000 service=storage-engine engine=tsm1 op_name=tsm1_cache_snapshot op_event=end op_elapsed=9352.816ms
ts=2026-08-14T12:47:07.855028Z lvl=info msg="TSM compaction (start)" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 op_event=start
ts=2026-08-14T12:47:07.856015Z lvl=info msg="Beginning compaction" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_files_n=8
ts=2026-08-14T12:47:07.856047Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=0 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000033-000000001.tsm
ts=2026-08-14T12:47:07.856073Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=1 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000034-000000001.tsm
ts=2026-08-14T12:47:07.856097Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=2 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000035-000000001.tsm
ts=2026-08-14T12:47:07.856119Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=3 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000036-000000001.tsm
ts=2026-08-14T12:47:07.856142Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=4 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000037-000000001.tsm
ts=2026-08-14T12:47:07.856164Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=5 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000038-000000001.tsm
ts=2026-08-14T12:47:07.856185Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=6 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000039-000000001.tsm
ts=2026-08-14T12:47:07.856206Z lvl=info msg="Compacting file" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 tsm1_index=7 tsm1_file=/var/lib/influxdb/engine/data/8e9af974e7714e8a/autogen/27676/000000041-000000001.tsm
ts=2026-08-14T12:47:14.400290Z lvl=info msg="TSM compaction (end)" log_id=14gzJF80000 service=storage-engine engine=tsm1 tsm1_level=1 tsm1_strategy=level op_name=tsm1_compact_group db_shard_id=27676 op_event=end op_elapsed=6545.275ms
panic: runtime error: slice bounds out of range [:1024] with capacity 1000
goroutine 1035540 [running]:
github.com/influxdata/influxdb/v2/tsdb/cursors.(*IntegerArray).Include(0x40019c53c8, 0x7fadd8ad38?, 0x18cb2bcecdde9b43)
        /root/project/tsdb/cursors/arrayvalues.gen.go:335 +0x1e4
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*tsmBatchKeyIterator).combineInteger(0x40233d0c60, 0x10?)
        /root/project/tsdb/engine/tsm1/compact.gen.go:320 +0x2b4
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*tsmBatchKeyIterator).mergeInteger(0x40233d0c60)
        /root/project/tsdb/engine/tsm1/compact.gen.go:263 +0x104
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*tsmBatchKeyIterator).merge(0x40233d0c60?)
        /root/project/tsdb/engine/tsm1/compact.go:1684 +0x34
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*tsmBatchKeyIterator).Next(0x40233d0c60)
        /root/project/tsdb/engine/tsm1/compact.go:1542 +0x80
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*Compactor).write(0x401d6fd000, {0x400b390420, 0x58}, {0x7fb1c429e0, 0x40233d0c60}, 0x1, 0x401d66a800)
        /root/project/tsdb/engine/tsm1/compact.go:1239 +0x298
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*Compactor).writeNewFiles(0x401d6fd000, 0x29, 0x1, {0x401d66a200?, 0x8, 0x8?}, {0x7fb1c429e0, 0x40233d0c60}, 0x1, 0x401d66a800)
        /root/project/tsdb/engine/tsm1/compact.go:1141 +0x278
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*Compactor).compact(0x401d6fd000, 0x0, {0x401d66a200, 0x8, 0x8}, 0x401d66a800, 0x3e8)
        /root/project/tsdb/engine/tsm1/compact.go:1031 +0x5c4
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*Compactor).CompactFull(0x401d6fd000, {0x401d66a200, 0x8, 0x8}, 0x401d66a800, 0x3e8)
        /root/project/tsdb/engine/tsm1/compact.go:1049 +0x194
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*compactionStrategy).compactGroup(0x40234d8620)
        /root/project/tsdb/engine/tsm1/engine.go:2565 +0x338
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*compactionStrategy).Apply(0x40234d8620)
        /root/project/tsdb/engine/tsm1/engine.go:2538 +0x34
github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*Engine).compactHiPriorityLevel.func1()
        /root/project/tsdb/engine/tsm1/engine.go:2407 +0x9c
created by github.com/influxdata/influxdb/v2/tsdb/engine/tsm1.(*Engine).compactHiPriorityLevel in goroutine 17214
        /root/project/tsdb/engine/tsm1/engine.go:2399 +0x16c

Aug 14 14:47:14 systemd[1]: influxdb.service: Main process exited, code=exited, status=2/INVALIDARGUMENT
Aug 14 14:47:14 systemd[1]: influxdb.service: Failed with result 'exit-code'.
Aug 14 14:47:14 systemd[1]: influxdb.service: Consumed 13min 41.047s CPU time.

And some system information:

pi@raspberrypi:~ $ cat /etc/os-release
PRETTY_NAME="Debian GNU/Linux 13 (trixie)"
NAME="Debian GNU/Linux"
VERSION_ID="13"
VERSION="13 (trixie)"
VERSION_CODENAME=trixie
DEBIAN_VERSION_FULL=13.6
ID=debian
HOME_URL="https://www.debian.org/"
SUPPORT_URL="https://www.debian.org/support"
BUG_REPORT_URL="https://bugs.debian.org/"
pi@raspberrypi:~ $ uname -a
Linux raspberrypi 6.18.29+rpt-rpi-v8 #1 SMP PREEMPT Debian 1:6.18.29-1+rpt1 (2026-05-12) aarch64 GNU/Linux
pi@raspberrypi:~ $ influxd version
InfluxDB v2.9.1 (git: d4fa1941fd) build_date: 2026-05-11T20:48:40Z

I’m currently upgrading the whole system (latest Kernel available for debian/raspian etc) but I don’t expect the issue to be solved by this as the golang dump clearly indicates this is an issue in the influx code (EDIT: yes it indeed turned out it did not change a thing)

Seems like it’s always right after the first (or maybe second) compaction process (indicated with TSM compaction (start) and TSM compaction (end))

It seems plausible that this is a bug in 2.9.1. If you restart the entire system with the file present, then restart the system again after removing the file, does it continue to happen regardless? If so, I’d recommend filing an issue in the InfluxDB repo with the 2.x tag.

Yes the only difference is that with removing the file, it took something like 2h until the next crash (I don’t write too much) while with keeping the file influx crashes already during startup (when loading the stored files/shards)

@atticus-sullivan This is stemming from corrupt TSM files as you have pointed out. Are you using the influx_inspect tool influxd inspect | InfluxDB OSS v2 Documentation to verify the TSM files on disk? I would suggest using this and removing all bad TSM files.

If you continue to see the error persist even after all corrupted files are dealt with:

Seeing as you’re running this on a raspberry pi I assume you are using an SD card for storage? There could be something wrong with the actual hardware itself since you’re seeing it consistently. Could you check for kernel IO errors? Maybe using dmesg like so:

dmesg | grep -iE 'mmc0|blk|i/o error'

specifically something like blk_update_request errors could be beneficial to see.

If you have another SD card it may be worth it to try out the same workload on there.

Save report to check what data will be missing

$ sudo influxd inspect report-tsm --data-path /var/lib/influxdb/engine/data > report.txt

analyze what files are corrupt and move them elsewhere

$ sudo influxd inspect verify-tsm --engine-path /var/lib/influxdb/engine
Broken Blocks: 27 / 7235913, in 703.767274329s
$ sudo influxd inspect verify-tsm --engine-path /var/lib/influxdb/engine -v 2>&1 | tee log
$ awk '!/healthy$/ {sub(/:$/, "", $1); print $1}' log > unhealthy
$ readarray -t array < unhealthy ; for ele in "${array[@]}" ; do sudo mv "$ele" ./ ; done

I’ll test tomorrow whether this actually helped as this takes time (in the past influx was fine for a couple of hours too)

And yes this is a raspberry pi, but I’m using an USB-Stick for storage and until now I can’t find any IO errors.

Runs for over 24 hours without crash, so it seems to be fixed by removing *all* corrupt files now