🦍 🦐 👡 Dipfake video one frame 🌨️ 👨🏾‍🔬 🎯

First Order Motion Model working example

Is it possible to make an entire movie from one photograph? And having recorded the movements of one person, replace him with another in the video? Of course, the answer to these questions is extremely important for areas such as cinema, photography, and the development of computer games. The solution could be digital photo processing using specialized software. The problem in question among specialists in this field is called the task of automatic synthesis of video or image animation.

To obtain the expected result, existing approaches combine objects extracted from the original image and movements that can be delivered as a separate video - “donor”.

Now, in most areas, image animation is done using computer graphics tools. This approach requires additional knowledge about the object that we want to animate - its 3D model is usually necessary (how it works now in the film industry can be found here ). Most of the latest solutions to this problem are based on in-depth training of models, which are based on generative-competitive neural networks (GAN) and variational autoencoders (VAE). These models usually use pre-trained modules to search for key points of objects in the image. The main problem with this approach is that these modules can only recognize the objects on which they were trained.

, ? «First Order Motion Model for Image Animation». — First Order Motion Model, . , (, , ), , .

…

, .

, , (occlusion map). . , , .

: .
$D \in \mathbb{R} ^{3×H×W}$ $S ∈ \mathbb{R} ^{3×H×W}$ . $S$ $D$ .

$S$ $D$ . , ( ) $R$ . $\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}$ $D$ $S$ $\hat{\mathcal{O}}_{\mathrm{S \leftarrow D}}$ . .

$\mathcal{T}_{\mathrm{S \leftarrow D}}$ $D$ $S$ . $\mathcal{T}_{\mathrm{S \leftarrow D}}$ . , $R$ ( ), $\mathcal{T}_{\mathrm{S \leftarrow D}}$ $\mathcal{T}_{\mathrm{S \leftarrow R}}$ $\mathcal{T}_{\mathrm{R \leftarrow D}}$ . , $X$ , $\mathcal{T}_{\mathrm{X \leftarrow R}}$ . $K$ $p_1,..., p_K$ , $p_1,..., p_K$ $R$ .

$\mathcal{T}_{\mathrm{R \leftarrow X}} = \mathcal{T}_{\mathrm{X \leftarrow R}}^{-1}$ , , $\mathcal{T}_{\mathrm{X \leftarrow R}}$ .

T_{S \leftarrow D} = T_{S \leftarrow R} \circ T_{R \leftarrow D} = T_{S \leftarrow R} \circ T_{D \leftarrow R}^{- 1}

$\mathcal{T}_{\mathrm{S \leftarrow D}} = \mathcal{T}_{\mathrm{S \leftarrow R}} \circ \mathcal{T}_{\mathrm{R \leftarrow D}} = \mathcal{T}_{\mathrm{S \leftarrow R}} \circ \mathcal{T}_{\mathrm{D \leftarrow R}}^{-1}$

$\mathcal{T}_{\mathrm{S \leftarrow R}}(p_k)$ $\mathcal{T}_{\mathrm{D \leftarrow R}}(p_k)$ . U-Net, $K$ , .
softmax , .

$P$ $\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}$ $\mathcal{T}_{\mathrm{S \leftarrow D}}(z)$ ( $z$ ), $S$ . , $\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}$ , , $D$ , $S$ . $\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}$ , $K$ $S^0,...,S^k$ ( $S^0 = S$ ), $\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}$ . $S^1,...,S^k$ U-Net.
$\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}(z)$ :

$M_k$ — ( $M_0$ — ) $J_k$ :

, $S$ $\hat{D}$ . , . down-sampling $\xi \in \mathbb{R}^{H' \times W'}$ . $\xi$ c $\hat{\mathcal{T}}_{\mathrm{S \leftarrow D}}$ . $S$ , $\hat{D}$ . — $\hat{\mathcal{O}}_{\mathrm{S \leftarrow D}} \in [0, 1]^{H' \times W'}$ , , , $S$ . :

ξ^{'} = {\hat{O}}_{S \leftarrow D} ⊙ f_{w} (ξ, {\hat{T}}_{S \leftarrow D})

$\xi ' = \hat{\mathcal{O}}_{\mathrm{S \leftarrow D}} \odot f_w(\xi, \hat{\mathcal{T}}_{\mathrm{S \leftarrow D}})$

$f_w(\cdot, \cdot)$ , $\odot$ — ( ).

, . $\xi '$ , .

, . reconstruction loss, . - VGG-19. reconstruction loss :

L_{r e c} (\hat{D}, D) = \sum_{i = 1}^{I} | N_{i} (\hat{D}) - N_{i} (D) |

$L_{rec} (\hat{D}, D)= \sum_{i = 1}^I |N_i(\hat{D}) - N_i(D)|$

$\hat{D}$ — , $D$ — , $N_i(\cdot)$ — i- , VGG-19, $I$ — .

- . . , . , . , , , .

, $X$ $\mathcal{T}_{\mathrm{X \leftarrow Y}}$ , , thin plane spline. $Y$ . , $\mathcal{T}_{\mathrm{X \leftarrow R}}$
$\mathcal{T}_{\mathrm{Y \leftarrow R}}$ . C :