MiniMax H3 Ref2VA + audio recasts identity, while FL2VA + Ref2V Turbo preserves it β€” expected behavior?

#91
by lorentaken - opened

I am testing MiniMax H3 with one human identity reference and a short speech-audio reference.

The same visual reference preserves identity reasonably well through ImageToVideo/FL2VA, but identity collapses when the reference is supplied through ReferenceToVideo together with audio.

Common setup

  • Resolution: 1344x768
  • FPS: 24
  • Length: 124 frames / 5.17 seconds
  • Euler sampler
  • Beta scheduler
  • Same source image
  • Same structured reference prompt
  • InsightFace buffalo_l cosine similarity used for diagnostics

Test A β€” ImageToVideo, visual only

  • FL2VA checkpoint
  • Ref2V Turbo LoRA
  • ImageToVideo
  • 8 steps
  • No audio

Results:

  • Mean similarity: 0.8855
  • Minimum similarity: 0.8628

The identity is visually preserved.

Test B β€” ReferenceToVideo with audio, ref_image_size=match

  • Same checkpoint
  • Same Ref2V Turbo LoRA
  • ReferenceToVideo
  • One image reference
  • One audio reference
  • 8 steps

Results:

  • Mean similarity: 0.0884
  • Minimum similarity: 0.0328

The output is effectively a different person.

Test C β€” ReferenceToVideo with audio, ref_image_size=max

Only the image sizing mode was changed.

Results:

  • Mean similarity: 0.0902
  • Minimum similarity: 0.0527

Changing match to max does not recover the identity.

Test D β€” Native Ref2VA without LoRA

  • Native Ref2VA checkpoint
  • No LoRA
  • ReferenceToVideo
  • Audio reference
  • 18 steps
  • ref_image_size=max

Results:

  • Mean similarity: 0.0669
  • Minimum similarity: 0.0350

Using the native checkpoint and more steps still produces the wrong identity.

Observation

The failure appears to be specifically associated with audio conditioning through the ReferenceToVideo path.

ImageToVideo/FL2VA preserves the identity much better with the same visual reference, while ReferenceToVideo with audio recasts the subject.

The following did not solve the problem:

  • changing match to max;
  • removing the Turbo LoRA;
  • increasing the number of steps to 18;
  • using the native Ref2VA checkpoint.

Questions

  1. Is this a known limitation of the current audio-conditioned ReferenceToVideo path?
  2. Is the Ref2V Turbo LoRA intended for visual reference conditioning only?
  3. Is there a recommended workflow for combining:
    • a fixed first-frame identity;
    • a separate image reference;
    • audio or lip synchronization?
  4. Should audio be connected through the FL2VA/ImageToVideo path instead?
  5. Are there recommended audio-shift, VAE, attention or sampler settings for preserving identity?
  6. Could this be caused by an incorrect ComfyUI node connection?

I am not sure i understand, Ref2va preserves identity with or without audio. I have no issues with the identity preservation for ref2va as it is the best i have seen for any local model by far in keeping subjects.

This should do what you need i think
First/Last frames and refs usable with the normal ref2va model. Notes should tell you how to prompt it. Allows full lip syncing via latent noise masking the audio channel if needed.
https://civitai.red/models/2847150/minimax-simple

Sign up or log in to comment