The Experts below are selected from a list of 172551 Experts worldwide ranked by ideXlab platform

Christian Theobalt - One of the best experts on this subject based on the ideXlab platform.

  • DeepCap: Monocular Human Performance Capture Using Weak Supervision
    arXiv: Computer Vision and Pattern Recognition, 2020
    Co-Authors: Marc Habermann, Michael Zollhoefer, Gerard Pons-moll, Christian Theobalt
    Abstract:

    Human Performance Capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous Performance Capture approaches either required expensive multi-view setups or did not recover dense space-time coherent geometry with frame-to-frame correspondences. We propose a novel deep learning approach for monocular dense human Performance Capture. Our method is trained in a weakly supervised manner based on multi-view supervision completely removing the need for training data with 3D ground truth annotations. The network architecture is based on two separate networks that disentangle the task into a pose estimation and a non-rigid surface deformation step. Extensive qualitative and quantitative evaluations show that our approach outperforms the state of the art in terms of quality and robustness.

  • CVPR - DeepCap: Monocular Human Performance Capture Using Weak Supervision
    2020 IEEE CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
    Co-Authors: Marc Habermann, Gerard Pons-moll, Michael Zollhöfer, Christian Theobalt
    Abstract:

    Human Performance Capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous Performance Capture approaches either required expensive multi-view setups or did not recover dense space-time coherent geometry with frame-to-frame correspondences. We propose a novel deep learning approach for monocular dense human Performance Capture. Our method is trained in a weakly supervised manner based on multi-view supervision completely removing the need for training data with 3D ground truth annotations. The network architecture is based on two separate networks that disentangle the task into a pose estimation and a non-rigid surface deformation step. Extensive qualitative and quantitative evaluations show that our approach outperforms the state of the art in terms of quality and robustness.

  • EgoFace: Egocentric Face Performance Capture and Videorealistic Reenactment
    arXiv: Computer Vision and Pattern Recognition, 2019
    Co-Authors: Mohamed Elgharib, Hans-peter Seidel, Ayush Tewari, Hyeongwoo Kim, Wentao Liu, Christian Theobalt
    Abstract:

    Face Performance Capture and reenactment techniques use multiple cameras and sensors, positioned at a distance from the face or mounted on heavy wearable devices. This limits their applications in mobile and outdoor environments. We present EgoFace, a radically new lightweight setup for face Performance Capture and front-view videorealistic reenactment using a single egocentric RGB camera. Our lightweight setup allows operations in uncontrolled environments, and lends itself to telepresence applications such as video-conferencing from dynamic environments. The input image is projected into a low dimensional latent space of the facial expression parameters. Through careful adversarial training of the parameter-space synthetic rendering, a videorealistic animation is produced. Our problem is challenging as the human visual system is sensitive to the smallest face irregularities that could occur in the final results. This sensitivity is even stronger for video results. Our solution is trained in a pre-processing stage, through a supervised manner without manual annotations. EgoFace Captures a wide variety of facial expressions, including mouth movements and asymmetrical expressions. It works under varying illuminations, background, movements, handles people from different ethnicities and can operate in real time.

  • livecap real time human Performance Capture from monocular video
    ACM Transactions on Graphics, 2019
    Co-Authors: Marc Habermann, Michael Zollhöfer, Weipeng Xu, Gerard Ponsmoll, Christian Theobalt
    Abstract:

    We present the first real-time human Performance Capture approach that reconstructs dense, space-time coherent deforming geometry of entire humans in general everyday clothing from just a single RGB video. We propose a novel two-stage analysis-by-synthesis optimization whose formulation and implementation are designed for high Performance. In the first stage, a skinned template model is jointly fitted to background subtracted input video, 2D and 3D skeleton joint positions found using a deep neural network, and a set of sparse facial landmark detections. In the second stage, dense non-rigid 3D deformations of skin and even loose apparel are Captured based on a novel real-time capable algorithm for non-rigid tracking using dense photometric and silhouette constraints. Our novel energy formulation leverages automatically identified material regions on the template to model the differing non-rigid deformation behavior of skin and apparel. The two resulting non-linear optimization problems per frame are solved with specially tailored data-parallel Gauss-Newton solvers. To achieve real-time Performance of over 25Hz, we design a pipelined parallel architecture using the CPU and two commodity GPUs. Our method is the first real-time monocular approach for full-body Performance Capture. Our method yields comparable accuracy with off-line Performance Capture techniques while being orders of magnitude faster.

  • LiveCap: Real-time Human Performance Capture from Monocular Video
    arXiv: Computer Vision and Pattern Recognition, 2018
    Co-Authors: Marc Habermann, Michael Zollhoefer, Gerard Pons-moll, Christian Theobalt
    Abstract:

    We present the first real-time human Performance Capture approach that reconstructs dense, space-time coherent deforming geometry of entire humans in general everyday clothing from just a single RGB video. We propose a novel two-stage analysis-by-synthesis optimization whose formulation and implementation are designed for high Performance. In the first stage, a skinned template model is jointly fitted to background subtracted input video, 2D and 3D skeleton joint positions found using a deep neural network, and a set of sparse facial landmark detections. In the second stage, dense non-rigid 3D deformations of skin and even loose apparel are Captured based on a novel real-time capable algorithm for non-rigid tracking using dense photometric and silhouette constraints. Our novel energy formulation leverages automatically identified material regions on the template to model the differing non-rigid deformation behavior of skin and apparel. The two resulting non-linear optimization problems per-frame are solved with specially-tailored data-parallel Gauss-Newton solvers. In order to achieve real-time Performance of over 25Hz, we design a pipelined parallel architecture using the CPU and two commodity GPUs. Our method is the first real-time monocular approach for full-body Performance Capture. Our method yields comparable accuracy with off-line Performance Capture techniques, while being orders of magnitude faster.

Yebin Liu - One of the best experts on this subject based on the ideXlab platform.

  • MulayCap: Multi-layer Human Performance Capture Using A Monocular Video Camera.
    IEEE Transactions on Visualization and Computer Graphics, 2020
    Co-Authors: Su Zhaoqi, Lu Fang, Weilin Wan, Lingjie Liu, Wenping Wang, Yebin Liu
    Abstract:

    We introduce MulayCap, a novel human Performance Capture method using a monocular video camera without the need for pre-scanning. The method uses "multi-layer" representations for geometry reconstruction and texture rendering, respectively. For geometry reconstruction, we decompose the clothed human into multiple geometry layers, namely a body mesh layer and a garment piece layer. The key technique behind is a Garment-from-Video (GfV) method for optimizing the garment shape and reconstructing the dynamic cloth to fit the input video sequence, based on a cloth simulation model effectively solved with gradient descent. For texture rendering, we decompose each input image frame into a shading layer and an albedo layer, and propose a method for fusing an albedo map and solving for detailed garment geometry using the shading layer. Compared with existing single view human Performance Capture systems, our "multi-layer" approach bypasses the tedious and time consuming scanning step for obtaining a human specific mesh template. Experimental results demonstrate that MulayCap produces realistic rendering of dynamically changing details that has not been achieved in any previous monocular video camera systems. Benefiting from its fully semantic modeling, MulayCap can be applied to various important editing applications, such as cloth editing, re-targeting, relighting, and AR applications.

  • SimulCap : Single-View Human Performance Capture with Cloth Simulation
    arXiv: Computer Vision and Pattern Recognition, 2019
    Co-Authors: Zerong Zheng, Gerard Pons-moll, Yuan Zhong, Jianhui Zhao, Qionghai Dai, Yebin Liu
    Abstract:

    This paper proposes a new method for live free-viewpoint human Performance Capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body, and (ii) a physics-based Performance Capture procedure. We first digitize the performer using multi-layer surface representation, which includes the undressed body surface and separate clothing meshes. For Performance Capture, we perform skeleton tracking, cloth simulation, and iterative depth fitting sequentially for the incoming frame. By incorporating cloth simulation into the Performance Capture pipeline, we can simulate plausible cloth dynamics and cloth-body interactions even in the occluded regions, which was not possible in previous Capture methods. Moreover, by formulating depth fitting as a physical process, our system produces cloth tracking results consistent with the depth observation while still maintaining physical constraints. Results and evaluations show the effectiveness of our method. Our method also enables new types of applications such as cloth retargeting, free-viewpoint video rendering and animations.

  • CVPR - SimulCap : Single-View Human Performance Capture With Cloth Simulation
    2019 IEEE CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
    Co-Authors: Zerong Zheng, Gerard Pons-moll, Yuan Zhong, Jianhui Zhao, Qionghai Dai, Yebin Liu
    Abstract:

    This paper proposes a new method for live free-viewpoint human Performance Capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body, and (ii) a physics-based Performance Capture procedure. We first digitize the performer using multi-layer surface representation, which includes the undressed body surface and separate clothing meshes. For Performance Capture, we perform skeleton tracking, cloth simulation, and iterative depth fitting sequentially for the incoming frame. By incorporating cloth simulation into the Performance Capture pipeline, we can simulate plausible cloth dynamics and cloth-body interactions even in the occluded regions, which was not possible in previous Capture methods. Moreover, by formulating depth fitting as a physical process, our system produces cloth tracking results consistent with the depth observation while still maintaining physical constraints. Results and evaluations show the effectiveness of our method. Our method also enables new types of applications such as cloth retargeting, free-viewpoint video rendering and animations.

  • ECCV (9) - HybridFusion: Real-Time Performance Capture Using a Single Depth Sensor and Sparse IMUs
    Computer Vision – ECCV 2018, 2018
    Co-Authors: Zerong Zheng, Qionghai Dai, Lu Fang, Kaiwen Guo, Yebin Liu
    Abstract:

    We propose a light-weight yet highly robust method for real-time human Performance Capture based on a single depth camera and sparse inertial measurement units (IMUs). Our method combines non-rigid surface tracking and volumetric fusion to simultaneously reconstruct challenging motions, detailed geometries and the inner human body of a clothed subject. The proposed hybrid motion tracking algorithm and efficient per-frame sensor calibration technique enable non-rigid surface reconstruction for fast motions and challenging poses with severe occlusions. Significant fusion artifacts are reduced using a new confidence measurement for our adaptive TSDF-based fusion. The above contributions are mutually beneficial in our reconstruction system, which enable practical human Performance Capture that is real-time, robust, low-cost and easy to deploy. Experiments show that extremely challenging Performances and loop closure problems can be handled successfully.

  • Human Performance Capture Using Multiple Handheld Kinects
    Computer Vision and Machine Learning with RGB-D Sensors, 2014
    Co-Authors: Yebin Liu, Qionghai Dai, Yangang Wang, Christian Theobalt
    Abstract:

    Capturing real Performances of human actors has been an important topic in the fields of computer graphics and computer vision in the last few decades. The reconstructed 3D Performance can be used for character animation and free-viewpoint video. While most of the available Performance Capture approaches rely on a 3D video studio with tens of RGB cameras, this chapter presents a method for marker-less Performance Capture of single or multiple human characters using only three handheld Kinects. Compared with the RGB camera approaches, the proposed method is more convenient with respect to data acquisition, allowing for much fewer cameras and carry-on camera Capture. The method introduced in this chapter reconstructs human skeletal poses, deforming surface geometry and camera poses for every time step of the depth video. It succeeds on general uncontrolled indoor scenes with potentially dynamic background, and it succeeds even for reconstruction of multiple closely interacting characters.

Vladimir Tankovich - One of the best experts on this subject based on the ideXlab platform.

  • Motion2fusion: real-time volumetric Performance Capture
    ACM Transactions on Graphics, 2017
    Co-Authors: Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Shahram Izadi
    Abstract:

    We present Motion2Fusion, a state-of-the-art 360 Performance Capture system that enables *real-time* reconstruction of arbitrary non-rigid scenes. We provide three major contributions over prior work: 1) a new non-rigid fusion pipeline allowing for far more faithful reconstruction of high frequency geometric details, avoiding the over-smoothing and visual artifacts observed previously. 2) a high speed pipeline coupled with a machine learning technique for 3D correspondence field estimation reducing tracking errors and artifacts that are attributed to fast motions. 3) a backward and forward non-rigid alignment strategy that more robustly deals with topology changes but is still free from scene priors. Our novel Performance Capture system demonstrates real-time results nearing 3x speed-up from previous state-of-the-art work on the exact same GPU hardware. Extensive quantitative and qualitative comparisons show more precise geometric and texturing results with less artifacts due to fast motions or topology changes than prior art.

  • fusion4d real time Performance Capture of challenging scenes
    International Conference on Computer Graphics and Interactive Techniques, 2016
    Co-Authors: Sameh Khamis, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Christoph Rhemann, Yury Degtyarev, Sergio Orts Escolano, Jonathan Taylor, Pushmeet Kohli, Vladimir Tankovich
    Abstract:

    We contribute a new pipeline for live multi-view Performance Capture, generating temporally coherent high-quality reconstructions in real-time. Our algorithm supports both incremental reconstruction, improving the surface estimation over time, as well as parameterizing the nonrigid scene motion. Our approach is highly robust to both large frame-to-frame motion and topology changes, allowing us to reconstruct extremely challenging scenes. We demonstrate advantages over related real-time techniques that either deform an online generated template or continually fuse depth data nonrigidly into a single reference model. Finally, we show geometric reconstruction results on par with offline methods which require orders of magnitude more processing time and many more RGBD cameras.

Edilson De Aguiar - One of the best experts on this subject based on the ideXlab platform.

  • Multi-view Performance Capture of Surface Details
    International Journal of Computer Vision, 2017
    Co-Authors: Nadia Robertini, Edilson De Aguiar, Dan Casas, Christian Theobalt
    Abstract:

    This paper presents a novel approach to recover true fine surface detail of deforming meshes reconstructed from multi-view video. Template-based methods for Performance Capture usually produce a coarse-to-medium scale detail 4D surface reconstruction which does not contain the real high-frequency geometric detail present in the original video footage. Fine scale deformation is often incorporated in a second pass by using stereo constraints, features, or shading-based refinement. In this paper, we propose an alternative solution to this second stage by formulating dense dynamic surface reconstruction as a global optimization problem of the densely deforming surface. Our main contribution is an implicit representation of a deformable mesh that uses a set of Gaussian functions on the surface to represent the initial coarse mesh, and a set of Gaussians for the images to represent the original Captured multi-view images. We effectively find the fine scale deformations for all mesh vertices, which maximize photo-temporal-consistency, by densely optimizing our model-to-image consistency energy on all vertex positions. Our formulation yields a smooth closed form energy with implicit occlusion handling and analytic derivatives. Furthermore, it does not require error-prone correspondence finding or discrete sampling of surface displacement values. We demonstrate our approach on a variety of datasets of human subjects wearing loose clothing and performing different motions. We qualitatively and quantitatively demonstrate that our technique successfully reproduces finer detail than the input baseline geometry.

  • efficient multi view Performance Capture of fine scale surface detail
    International Conference on 3D Vision, 2014
    Co-Authors: Nadia Robertini, Edilson De Aguiar, Thomas Helten, Christian Theobalt
    Abstract:

    We present a new effective way for Performance Capture of deforming meshes with fine-scale time-varying surface detail from multi-view video. Our method builds up on coarse 4D surface reconstructions, as obtained with commonly used template-based methods. As they only Capture models of coarse-to-medium scale detail, fine scale deformation detail is often done in a second pass by using stereo constraints, features, or shading-based refinement. In this paper, we propose a new effective and stable solution to this second step. Our framework creates an implicit representation of the deformable mesh using a dense collection of 3D Gaussian functions on the surface, and a set of 2D Gaussians for the images. The fine scale deformation of all mesh vertices that maximizes photo-consistency can be efficiently found by densely optimizing a new model-to-image consistency energy on all vertex positions. A principal advantage is that our problem formulation yields a smooth closed form energy with implicit occlusion handling and analytic derivatives. Error-prone correspondence finding, or discrete sampling of surface displacement values are also not needed. We show several reconstructions of human subjects wearing loose clothing, and we qualitatively and quantitatively show that we robustly Capture more detail than related methods.

  • 3DV - Efficient Multi-view Performance Capture of Fine-Scale Surface Detail
    2014 2nd International Conference on 3D Vision, 2014
    Co-Authors: Nadia Robertini, Edilson De Aguiar, Thomas Helten, Christian Theobalt
    Abstract:

    We present a new effective way for Performance Capture of deforming meshes with fine-scale time-varying surface detail from multi-view video. Our method builds up on coarse 4D surface reconstructions, as obtained with commonly used template-based methods. As they only Capture models of coarse-to-medium scale detail, fine scale deformation detail is often done in a second pass by using stereo constraints, features, or shading-based refinement. In this paper, we propose a new effective and stable solution to this second step. Our framework creates an implicit representation of the deformable mesh using a dense collection of 3D Gaussian functions on the surface, and a set of 2D Gaussians for the images. The fine scale deformation of all mesh vertices that maximizes photo-consistency can be efficiently found by densely optimizing a new model-to-image consistency energy on all vertex positions. A principal advantage is that our problem formulation yields a smooth closed form energy with implicit occlusion handling and analytic derivatives. Error-prone correspondence finding, or discrete sampling of surface displacement values are also not needed. We show several reconstructions of human subjects wearing loose clothing, and we qualitatively and quantitatively show that we robustly Capture more detail than related methods.

  • Image and Geometry Processing for 3-D Cinematography - Performance Capture from Multi-View Video
    Geometry and Computing, 2010
    Co-Authors: Christian Theobalt, Hans-peter Seidel, Edilson De Aguiar, Carsten Stoll, Sebastian Thrun
    Abstract:

    Nowadays, increasing Performance of computing hardware makes it feasible to simulate ever more realistic humans even in real-time applications for the end-user. To fully capitalize on these computational resources, all aspects of the human, including textural appearance and lighting, and, most importantly, dynamic shape and motion have to be simulated at high fidelity in order to convey the impression of a realistic human being. In consequence, the increase in computing power is flanked by increasing requirements to the skills of the animators. In this chapter, we describe several recently developed Performance Capture techniques that enable animators to measure detailed animations from real world subjects recorded on multi-view video. In contrast to classical motion Capture, Performance Capture approaches don’t only measure motion parameters without the use of optical markers, but also measure detailed spatio-temporally coherent dynamic geometry and surface texture of a performing subject. This chapter gives an overview of recent state-of-the-art Performance Capture approaches from the literature. The core of the chapter describes a new mesh-based Performance Capture algorithm that uses a combination of deformable surface and volume models for high-quality reconstruction of people in general apparel, i.e. also wide dresses and skirts. The chapter concludes with a discussion of the different approaches, pointers to additional literature and a brief outline of open research questions for the future.

  • Performance Capture from multi view video
    Image and Geometry Processing for 3-D Cinematography, 2010
    Co-Authors: Christian Theobalt, Hans-peter Seidel, Edilson De Aguiar, Carsten Stoll, Sebastian Thrun
    Abstract:

    Nowadays, increasing Performance of computing hardware makes it feasible to simulate ever more realistic humans even in real-time applications for the end-user. To fully capitalize on these computational resources, all aspects of the human, including textural appearance and lighting, and, most importantly, dynamic shape and motion have to be simulated at high fidelity in order to convey the impression of a realistic human being. In consequence, the increase in computing power is flanked by increasing requirements to the skills of the animators. In this chapter, we describe several recently developed Performance Capture techniques that enable animators to measure detailed animations from real world subjects recorded on multi-view video. In contrast to classical motion Capture, Performance Capture approaches don’t only measure motion parameters without the use of optical markers, but also measure detailed spatio-temporally coherent dynamic geometry and surface texture of a performing subject. This chapter gives an overview of recent state-of-the-art Performance Capture approaches from the literature. The core of the chapter describes a new mesh-based Performance Capture algorithm that uses a combination of deformable surface and volume models for high-quality reconstruction of people in general apparel, i.e. also wide dresses and skirts. The chapter concludes with a discussion of the different approaches, pointers to additional literature and a brief outline of open research questions for the future.

Thabo Beeler - One of the best experts on this subject based on the ideXlab platform.

  • Accurate markerless jaw tracking for facial Performance Capture
    ACM Transactions on Graphics, 2019
    Co-Authors: Gaspard Zoss, Thabo Beeler, Markus Gross, Derek Bradley
    Abstract:

    We present the first method to accurately track the invisible jaw based solely on the visible skin surface, without the need for any markers or augmentation of the actor. As such, the method can readily be integrated with off-the-shelf facial Performance Capture systems. The core idea is to learn a non-linear mapping from the skin deformation to the underlying jaw motion on a dataset where ground-truth jaw poses have been acquired, and then to retarget the mapping to new subjects. Solving for the jaw pose plays a central role in visual effects pipelines, since accurate jaw motion is required when retargeting to fantasy characters and for physical simulation. Currently, this task is performed mostly manually to achieve the desired level of accuracy, and the presented method has the potential to fully automate this labour intense and error prone process.

  • Real-time high-fidelity facial Performance Capture
    ACM Transactions on Graphics, 2015
    Co-Authors: Chen Cao, Derek Bradley, Kun Zhou, Thabo Beeler
    Abstract:

    We present the first real-time high-fidelity facial Capture method. The core idea is to enhance a global real-time face tracker, which provides a low-resolution face mesh, with local regressors that add in medium-scale details, such as expression wrinkles. Our main observation is that although wrinkles appear in different scales and at different locations on the face, they are locally very self-similar and their visual appearance is a direct consequence of their local shape. We therefore train local regressors from high-resolution Capture data in order to predict the local geometry from local appearance at runtime. We propose an automatic way to detect and align the local patches required to train the regressors and run them efficiently in real-time. Our formulation is particularly designed to enhance the low-resolution global tracker with exactly the missing expression frequencies, avoiding superimposing spatial frequencies in the result. Our system is generic and can be applied to any real-time tracker that uses a global prior, e.g. blend-shapes. Once trained, our online Capture approach can be applied to any new user without additional training, resulting in high-fidelity facial Performance reconstruction with person-specific wrinkle details from a monocular video camera in real-time.

  • High-quality passive facial Performance Capture using anchor frames
    ACM Transactions on Graphics, 2011
    Co-Authors: Thabo Beeler, Derek Bradley, Fabian Hahn, Bernd Bickel, Paul Beardsley, Craig Gotsman, Robert W. Sumner, Markus Gross
    Abstract:

    We present a new technique for passive and markerless facial Performance Capture based on anchor frames. Our method starts with high resolution per-frame geometry acquisition using state-of-the-art stereo reconstruction, and proceeds to establish a single triangle mesh that is propagated through the entire Performance. Leveraging the fact that facial Performances often contain repetitive subsequences, we identify anchor frames as those which contain similar facial expressions to a manually chosen reference expression. Anchor frames are automatically computed over one or even multiple Performances. We introduce a robust image-space tracking method that computes pixel matches directly from the reference frame to all anchor frames, and thereby to the remaining frames in the sequence via sequential matching. This allows us to propagate one reconstructed frame to an entire sequence in parallel, in contrast to previous sequential methods. Our anchored reconstruction approach also limits tracker drift and robustly handles occlusions and motion blur. The parallel tracking and mesh propagation offer low computation times. Our technique will even automatically match anchor frames across different sequences Captured on different occasions, propagating a single mesh to all Performances.