r/qnap
Jellyfin hardware transcoding fails with cuInit failed: CUDA_ERROR_NOT_INITIALIZED after some uptime
- upvotes
- 1
- comments
- 3
Post
Highlighted: the lines this signal was extracted from
I have a TS-855eU with 64 GB RAM, NVIDIA RTX A400 assigned to ContainerStation, QNAP NVIDIA GPU driver 575.64.05 through NvKernelDriver 5.2.10.3577. ContainerStation with linuxserver/jellyfin:12.1, NVENC transcoding runtime: nvidia-runtime environment: - NVIDIA_VISIBLE_DEVICES=all - NVIDIA_DRIVER_CAPABILITIES=compute,video,utility Any movie that needs transcoding fails at once. The Jellyfin log only shows FFmpeg exited with code 187. The FFmpeg transcode log shows: [CUDA] cu->cuInit(0) failed -> CUDA_ERROR_NOT_INITIALIZED: initialization error Device creation failed: -542398533. nvidia-smi inside the container sees the GPU, and the driver version matches the host/dev/nvidia0, nvidiactl, nvidia-uvm and nvidia-uvm-tools are all passed through and can be opened from inside the containerlibcuda.so in the container is the same version as the kernel module Hosts dmesg: NVRM: nvCheckOkFailedNoLog: Check failed: Out of memory [NV_ERR_NO_MEMORY] (0x00000051) returned from _memdescAllocInternal(pMemDesc) @ mem_desc.c:1353 NVRM: faultbufCtrlCmdMmuFaultBufferRegisterNonReplayBuf_IMPL: Error allocating client shadow fault buffer for non-replayable faults After some runtime the NAS only has about 700MB of free RAM. The rest is file cache. As a workaround i can run Qboosts Optimize Memory and the transcoding will work right away after that.Setting vm.min_free_kbytes to 4GB using sudo...
Keep reading with a free account
The rest of this post, and every signal for Nvidia, is in your free account.
Also quoted as evidence
According to https://bbs.archlinux.org/viewtopic.php?id=307380 this is a bug in the Driver, that was fixed with 580.76.05 and newer (at least for Arch).
From the post