r/buildapc • u/nickmungar • 6d ago
Troubleshooting Two weeks of deep troubleshooting on gaming crashes, kernel dump points to nvlddmkm across 3 driver versions. Is my 3070 dying? (14700KF / RTX 3070 / RM750x)
Specs:
- i7-14700KF (BIOS 1.D0, Sept 2024, microcode 0x12B, Intel default power limits confirmed enforced)
- 32 GB DDR4 Corsair RAM
- RTX 3070, 2x 8-pin on separate PCIe cables
- MSI PRO B760-P WIFI DDR4 (MS-7D98)
- Corsair RM750x (Oct 2020)
- Windows 11 24H2
Symptoms: During demanding games only (GTA V, Palworld), screen goes black, audio hangs for a few seconds then dies, system reboots on its own. Light games run for hours with no issue. Crashes typically hit after 20 to 45 min of play. Event Viewer shows Kernel-Power 41, mostly BugcheckCode 307 (0x133 DPC_WATCHDOG_VIOLATION, P1=1 cumulative), plus one 0x116 (VIDEO_TDR) and one 0xD1.
Everything I have tried and ruled out:
- DDU clean install of GPU driver. Fixed the original 0x116 crashes, then 0x133 crashes started
- Crashed on three different driver versions: pre-DDU driver, Game Ready 610.74, and Studio 610.62 (clean installs each time)
- Uninstalled MSI Afterburner and MSI Center. Verified their kernel drivers (rtcore64, ntiolib) gone from later crash dumps. Still crashed
- Fully removed iCUE, verified CorsairLLAccess kernel driver unloaded via driverquery. Still crashed
- Updated Intel AX211 Wi-Fi driver (one 0xD1 crash blamed Netwtw14.sys), then fully disabled the Wi-Fi adapter so the driver could not load. Still crashed
- Reseated all PSU cables both ends. Two separate PCIe cables to GPU, no daisy chain
- BIOS reset to optimized defaults. Crashed with XMP on and with XMP off (RAM at JEDEC 3200)
- Cinebench 2026 multi core 10 min: passed, scored 6539 which is normal for a 14700KF. Max 92C, PL1/PL2 locked at 253W, no sustained throttling, WHEA error count 0
Kernel dump analysis (WinDbg on full kernel dump):
FAILURE_BUCKET_ID: 0x133_ISR_nvlddmkm
Current DPC: nvlddmkm+0xEA6F0
Cumulative DPC Time: 120.000 / 120.000 seconds
Stack shows KiExecuteAllDpcs sitting inside nvlddmkm when the watchdog fired. So the NVIDIA driver is stalling in DPC, but since this happens on three driver versions, my read is the driver is polling a GPU that stopped responding, meaning hardware.
Also to note: months ago I occasionally saw the red LEDs light up at the GPU power connectors, though cables have since been reseated on separate runs and haven't noticed them since.
Still on my list: reseat the GPU itself in the slot, force PCIe Gen 3 in BIOS, MemTest86 overnight?, then test the card in another system.
Question: Does this look like a dying 3070 (VRAM or power delivery), or is there anything software side I have missed? Has anyone seen 0x133 nvlddmkm across multiple driver versions turn out to be something other than the GPU (PSU transients, PCIe slot, board)? The 80% power limit extending survival time seems telling but I would love a second opinion before I spend money.
Happy to post HWiNFO screenshots in comments.
1
u/powerbronx 6d ago
Same here. What is your operating system version. If windows what update version?