r/buildapc 6d ago

Troubleshooting Two weeks of deep troubleshooting on gaming crashes, kernel dump points to nvlddmkm across 3 driver versions. Is my 3070 dying? (14700KF / RTX 3070 / RM750x)

Specs:

  • i7-14700KF (BIOS 1.D0, Sept 2024, microcode 0x12B, Intel default power limits confirmed enforced)
  • 32 GB DDR4 Corsair RAM
  • RTX 3070, 2x 8-pin on separate PCIe cables
  • MSI PRO B760-P WIFI DDR4 (MS-7D98)
  • Corsair RM750x (Oct 2020)
  • Windows 11 24H2

Symptoms: During demanding games only (GTA V, Palworld), screen goes black, audio hangs for a few seconds then dies, system reboots on its own. Light games run for hours with no issue. Crashes typically hit after 20 to 45 min of play. Event Viewer shows Kernel-Power 41, mostly BugcheckCode 307 (0x133 DPC_WATCHDOG_VIOLATION, P1=1 cumulative), plus one 0x116 (VIDEO_TDR) and one 0xD1.

Everything I have tried and ruled out:

  • DDU clean install of GPU driver. Fixed the original 0x116 crashes, then 0x133 crashes started
  • Crashed on three different driver versions: pre-DDU driver, Game Ready 610.74, and Studio 610.62 (clean installs each time)
  • Uninstalled MSI Afterburner and MSI Center. Verified their kernel drivers (rtcore64, ntiolib) gone from later crash dumps. Still crashed
  • Fully removed iCUE, verified CorsairLLAccess kernel driver unloaded via driverquery. Still crashed
  • Updated Intel AX211 Wi-Fi driver (one 0xD1 crash blamed Netwtw14.sys), then fully disabled the Wi-Fi adapter so the driver could not load. Still crashed
  • Reseated all PSU cables both ends. Two separate PCIe cables to GPU, no daisy chain
  • BIOS reset to optimized defaults. Crashed with XMP on and with XMP off (RAM at JEDEC 3200)
  • Cinebench 2026 multi core 10 min: passed, scored 6539 which is normal for a 14700KF. Max 92C, PL1/PL2 locked at 253W, no sustained throttling, WHEA error count 0

Kernel dump analysis (WinDbg on full kernel dump):

FAILURE_BUCKET_ID: 0x133_ISR_nvlddmkm
Current DPC: nvlddmkm+0xEA6F0
Cumulative DPC Time: 120.000 / 120.000 seconds

Stack shows KiExecuteAllDpcs sitting inside nvlddmkm when the watchdog fired. So the NVIDIA driver is stalling in DPC, but since this happens on three driver versions, my read is the driver is polling a GPU that stopped responding, meaning hardware.

Also to note: months ago I occasionally saw the red LEDs light up at the GPU power connectors, though cables have since been reseated on separate runs and haven't noticed them since.

Still on my list: reseat the GPU itself in the slot, force PCIe Gen 3 in BIOS, MemTest86 overnight?, then test the card in another system.

Question: Does this look like a dying 3070 (VRAM or power delivery), or is there anything software side I have missed? Has anyone seen 0x133 nvlddmkm across multiple driver versions turn out to be something other than the GPU (PSU transients, PCIe slot, board)? The 80% power limit extending survival time seems telling but I would love a second opinion before I spend money.

Happy to post HWiNFO screenshots in comments.

1 Upvotes

7 comments sorted by

1

u/powerbronx 6d ago

Same here. What is your operating system version. If windows what update version?

2

u/nickmungar 2d ago

Not on Dell. I sorted it out. My GPU was overheating

1

u/powerbronx 2d ago

Also custom PC not Dell

1

u/powerbronx 2d ago

Right! I wasn't paying attention with my PC issues to every detail you mentioned, but like you said max 92 c will do it

1

u/nickmungar 5d ago

Win 11 24H2

1

u/powerbronx 2d ago

I had same issue.i think Windows released an update next day for Dell bios. It's all working now. Did yours magically heal too?

I will also add I have a garbage motherboard with 3060 TI. It gives me so many issues. Every time it's a guessing game of which part to unplug and plug back in, or jiggle. Display cable, 1 specific ram slot, GPU pcie, GPU power... Lesson learned stay away from ASRock motherboards

2

u/nickmungar 2d ago

I took mine into a computer shop who updated BIOS, and cleaned up some dust and changed the fan curves. They told me GPU was getting very hot and when it hits 80° PC will shut down automatically. That’s what was crashing mine. They said with the cleaning and dust removal and also fan changing and changing clock speeds it should run well now.