Rank 5: Hypervisor
So recently, after the clickhouse volume issue, i saw there was a lot of failed stats jobs, so i retried them with : queue-retry --name=v1-stats-usage --limit=40000 (resources too)
and saw there was a bump in the redis objects (reaching ~120k) and even after all job where completed nothing went down.
After some check it seems that the queue-retry recreate a new key with a new pid, and don't clear the old one, leaving orphan pids somewhere in the redis database.
After cleanup with AI, it went down to ~35k objects, and all stats where added back (even tho it didn't have the right time for them)