-
Notifications
You must be signed in to change notification settings - Fork 105
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#1576 In NVIDIA/NVSentinel;
[Bug]: GPU health monitor can lose connectivity failure before DCGM cleanup
priority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.[Bug]: metadata-collector exits 1 when patching a Pod that has already been deleted (NotFound treated as fatal)
bugSomething isn't workingSomething isn't workingpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1568 In NVIDIA/NVSentinel;[Bug]: Driver-pod detection in log-collector and labeler breaks under gpu-operator's nvidiaDriverCRD: true mode
bugSomething isn't workingSomething isn't workingpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.labeler is continuously OOMKilled (CrashLoopBackOff) at a 256Mi memory limit under a 5,000-node fleet
documentationImprovements or additions to documentationImprovements or additions to documentationpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1563 In NVIDIA/NVSentinel;MongoDB replicaSet loses primary under the platform connector pool load
documentationImprovements or additions to documentationImprovements or additions to documentationpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1562 In NVIDIA/NVSentinel;Microbenchmark labeler
priority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1561 In NVIDIA/NVSentinel;[Bug]: MongoDB inherits the host nofile soft limit which may lead to crashes at high node scale
bugSomething isn't workingSomething isn't workingpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1555 In NVIDIA/NVSentinel;[Feature]: Support recovery events for health-events-analyzer derived conditions
enhancementNew feature or requestNew feature or requestpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1553 In NVIDIA/NVSentinel;[Feature]: Enforce a maximum number of remediation attempts per node
enhancementNew feature or requestNew feature or requestpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1543 In NVIDIA/NVSentinel;[Bug]: Missed delete events will prevent kubernetes-object-monitor recovery
bugSomething isn't workingSomething isn't workingpriority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.[Docs]: Add documentation for all modules
priority/P1Max fix SLA: 183 daysMax fix SLA: 183 daysStatus: Open.#1527 In NVIDIA/NVSentinel;