Makes the common Linux H.264 export fully GPU-resident: decode (HW NV12) →
composite → VAAPI encode all on one device, no CPU round-trips.
- gpu-video-encoder: ZeroCopyEncoder now holds wgpu (device,queue,adapter) handles
instead of owning a DrmDevice. New `new_on_device(device,queue,adapter,...)` runs
the RGBA→NV12 render + DMA-BUF import on a passed device; `new` keeps building its
own. The encoder only *imports* VAAPI surfaces (not export), so the shared
import-capable device works directly.
- editor: stash the shared device handles on EditorApp (set in the creation closure
from the eframe render_state when the shared device is active) and thread them
through start_video_export / start_video_with_audio_export → try_build_zero_copy,
which uses new_on_device when available. ZeroCopyVideo carries on_shared_device →
the export composite uses hardware_ok=true (consuming HW-decoded GPU frames from
pt 1) only then.
Falls back to the own-device encoder (decode downloads to CPU) when no shared
device. Tradeoff: the export now shares the GPU with the UI thread (was a separate
device); acceptable for the GPU-resident-decode win. Compiles; crate tests build.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
10/12-bit HDR decodes to P010-style VAAPI surfaces (16-bit planes), which the
8-bit NV12 import couldn't handle — real HDR clips fell back to software (losing
the HDR). Import them as R16/Rg16 plane textures:
- dmabuf::Nv12DmaBuf gains `ten_bit`; import_raw picks R16_UNORM/R16G16_UNORM (vk)
+ R16Unorm/Rg16Unorm (wgpu) when set, else the existing R8/Rg8.
- hw_video.rs detects 10-bit from the hw frames context sw_format (P010/P012/P016)
and sets it. The NV12→RGB shader is unchanged: it samples normalized floats, and
the limited/full de-quant lands at the same values regardless of bit depth.
- vk_device requests TEXTURE_FORMAT_16BIT_NORM when the adapter supports it (R16/
Rg16 need it); absent → 10-bit falls back to software, 8-bit unaffected.
Pairs with the PQ/HLG/BT.2020 colour math so HDR10/HLG clips now reach the GPU
path end-to-end. NV12 callers (encoder, decode primitive) pass ten_bit=false.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
import_raw now takes (&wgpu::Device, &wgpu::Adapter) and extracts the raw Vulkan
handles via as_hal, instead of taking the crate's own DrmDevice. This lets a
VAAPI surface be imported onto ANY DMA-BUF-import-capable wgpu device — the
encoder/decoder's own device today, and the editor's shared device (Stage 3a)
next, so hardware-decoded frames are usable by the preview compositor.
Encoder/decoder pass their DrmDevice's device+adapter; round-trip decode test
still passes.
The zero-copy VAAPI encoder emitted full-range BT.709 NV12 but wrote no color
tags, so players assumed limited range and stretched it — the H.264 output
looked dark and oversaturated (preview and VP9/software were fine).
- Rgba2Nv12 takes `full_range` and applies the matching Y/chroma scale+offset
(limited 16-235 / full 0-255) via a uniform; the encoder sets color_range +
bt709 colorspace/primaries/transfer tags to match. ffprobe-confirmed.
- New ColorRange { Limited, Full } on VideoExportSettings (default Limited, the
universally-correct choice; serde(default) so saved settings still load),
surfaced as a "Color range" dropdown in advanced export settings for H.264.
The swscale software fallback still emits Limited regardless of the toggle
(Full only affects the VAAPI zero-copy path).
Three fixes found while running the zero-copy export on real Intel hardware:
- vaapi::create_device() retries LIBVA_DRIVER_NAME in order iHD -> auto ->
i965 -> radeonsi. libva was auto-selecting the legacy i965 driver, which
fails on newer Intel GPUs; the modern iHD (intel-media-driver) is needed.
encoder.rs now builds its hwdevice through this helper.
- vk_device: request the adapter's full limits instead of downlevel_defaults.
Vello's compute pipelines need max_storage_buffers_per_shader_stage >= 5
(downlevel caps at 4), which panicked Vello's shader init on the export
device. This device only ever runs on a real VAAPI GPU.
- ZeroCopyEncoder: unsafe impl Send. It owns its FFmpeg/Vulkan handles
exclusively and is only moved (onto the export thread), never shared.
ZeroCopyEncoder::new now takes an output path and writes a real container
(format inferred from the extension, e.g. .mp4): create an output format
context, add the h264 stream from the encoder, write header; encode_rgba
rescales each packet's ts and av_interleaved_write_frame's it; finish
flushes + writes the trailer + closes. Sets AV_CODEC_FLAG_GLOBAL_HEADER for
mp4/mov so SPS/PPS land in extradata. This lets the editor's existing
mux_video_and_audio consume the temp video file unchanged.
The zerocopy_encode test now writes a .mp4 and ffprobe-verifies the codec,
dimensions, and frame count. Also let wgpu own the imported plane-image
destruction via texture_from_raw drop callbacks (clears two warnings).
Build the full end-to-end zero-copy encoder, validated on Intel/VAAPI:
- render_nv12: fragment-shader RGBA->NV12 that renders luma/chroma into
the imported R8/RG8 plane render targets (compute storage can't write
the DMA-BUF-backed planes; render attachments can).
- dmabuf: import_raw imports an NV12 DMA-BUF by explicit layout; the two
plane images + shared memory are now destroyed by wgpu via texture_from_raw
drop callbacks (Arc MemoryGuard frees the memory once both images are
gone, in wgpu's wait-idle'd deferred pass) -- fixes the teardown segfault.
- encoder::ZeroCopyEncoder: renders an RGBA texture straight into a pooled
VAAPI surface (imports cached by VASurface id) and encodes with h264_vaapi.
encode_rgba + finish; the caller renders on device().
Tests: real-frame render into the surface matches the CPU NV12 reference,
and a 30-frame encode produces valid H.264 (ffprobe-verified) with clean
teardown. Not yet wired into the editor.