The Experts below are selected from a list of 66 Experts worldwide ranked by ideXlab platform

Alexander G. Hauptmann - One of the best experts on this subject based on the ideXlab platform.

  • real time near duplicate elimination for web video search with content and context
    IEEE Transactions on Multimedia, 2009
    Co-Authors: Chongwah Ngo, Alexander G. Hauptmann, Hungkhoon Tan
    Abstract:

    With the exponential growth of social media, there exist huge numbers of near-duplicate web videos, ranging from simple formatting to complex mixture of different editing effects. In addition to the abundant video content, the social Web provides rich sets of context information associated with web videos, such as Thumbnail Image, time duration and so on. At the same time, the popularity of Web 2.0 demands for timely response to user queries. To balance the speed and accuracy aspects, in this paper, we combine the contextual information from time duration, number of views, and Thumbnail Images with the content analysis derived from color and local points to achieve real-time near-duplicate elimination. The results of 24 popular queries retrieved from YouTube show that the proposed approach integrating content and context can reach real-time novelty re-ranking of web videos with extremely high efficiency, where the majority of duplicates can be rapidly detected and removed from the top rankings. The speedup of the proposed approach can reach 164 times faster than the effective hierarchical method proposed in , with just a slight loss of performance.

  • informedia at trecvid 2003 analyzing and searching broadcast news video
    TRECVID Workshop, 2003
    Co-Authors: Alexander G. Hauptmann, Michael G. Christel, Rong Jin, Norman Papernick, Pinar Duygulu, Chang Huang, Neema Moraveji, George Tzanetakis, J Yang, Rong Yan
    Abstract:

    Abstract : A concentrated effort was made by the authors to develop an interface allowing a human to succeed with video topics as defined in TRECVID 2001. This interface was part of the TRECVID 2002 interactive query task, in which a person could issue multiple queries and refinements to the video corpus in formulating the shot answer set for the topic at hand. The interface was designed to present a visually rich set of Thumbnail Images to the user, tailored for expert control over the number, scale, and attributes of the Images. Armed with this interface, an expert user completely familiar with the retrieval system and its features, but having no a priori knowledge of the TRECVID 2002 search test corpus, performed well on the search tasks. This exact system as used in the TRECVID 2002 interactive query task was again used for the TRECVID 2003 evaluation. To facilitate better visual browsing, we extended the storyboard idea to show keyframes across multiple video documents, where a document is automatically derived by segmenting a video production into story units through speech, silence, black frames, and other heuristics. The hierarchy of information units is frame, shot, document and full production. A set of documents is returned by a query. The shots for these documents are presented in a single storyboard, i.e., an ordered set of keyframes presented simultaneously on the computer screen, one keyframe per shot. Without further filtering, most queries would overwhelm the user with too many Images. Through the use of query context, the cardinality of the Image set can be greatly reduced. The search engine for text queries makes use of the Okapi method. The multiple document storyboard can be set to show only the shots containing matching words. This strategy of selecting a single Thumbnail Image to represent a video document based on query context resulted in more efficient information retrieval with greater user satisfaction.

  • video retrieval with the informedia digital video library system
    Text REtrieval Conference, 2002
    Co-Authors: Alexander G. Hauptmann, Rong Jin, Norman Papernick, Ricky Houghton, Sue Thornton
    Abstract:

    Background: The Informedia Digital Video Library System. The Informedia Digital Video Library [1] was the only NSF DLI project focusing specifically on information extraction from video and audio content. Over a terabyte of online data was collected, with automatically generated metadata and indices for retrieving videos from this library. The architecture for the project was based on the premise that real-time constraints on library and associated metadata creation could be relaxed in order to realize increased automation and deeper parsing and indexing for identifying the library contents and breaking it into segments. Library creation was an offline activity, with library exploration by users occurring online and making use of the generated metadata and segmentation. The goal of the Informedia interface was to enable quick access to relevant information in a digital video library, leveraging from derived metadata and the partitioning of the video into small segments. Figure 1 shows the IDVLS interface following a query. In this figure, a set of results is displayed at the bottom. The display includes a window containing a headline, and a pictorial menu of video segments each represented with a Thumbnail Image at approximately 1⁄4 resolution of the video in the horizontal and vertical dimensions. The headline window automatically pops up whenever the mouse is positioned over a result item; the headline window for the first result is shown. IDVLS also supports other ways of navigating and browsing the digital video library. These interface features were essential to deal with the ambiguity of the derived data generated by speech recognition, Image processing, and natural language processing. Consider the filmstrip and video playback IDVLS window shown in Figure 2. For this actual video in the IDVLS library, the segmentation process failed, resulting in a thirty-minute segment. This long segment was one of the returned results for the query “Mir collision.” The filmstrip in Figure 2 shows that the segment is more than just a story on the Russian space station, but rather begins with a commercial, then the weather, and then coverage of Hong Kong before addressing Mir. By overlaying the filmstrip and video playback windows with match location information, the user can quickly see that matches don’t occur until later in the segment, after these other stories that were irrelevant to the query. The match bars are optionally color-coded to specific query words; in Figure 2 “Mir” matches are in red and “collision” matches in purple. When the user moved the mouse over the match bars in the filmstrip, a text window displayed the actual matching word from the transcript or Video OCR metadata for that particular match; “Mir” is shown in one such text window in Figure 2. By investigating the distribution of match locations on the filmstrip, the user can determine the relevance of the returned result and the location of interest within the segment. The user can click on a match bar to jump directly to that point in the video segment. Hence, clicking the mouse as shown in Figure 2 would start playing the video at this mention of “Mir” with the overhead shot of people at desks. Similarly, IDVLS provided “seek to next match” and “seek to previous match” buttons in the video player allowing the user to quickly jump from one match to the next. In the example of Figure 2, these interface features allowed the user to bypass problems in segmentation and jump directly to the “Mir” story without having to first watch the opening video on other topics.

Stina Teilmannlock - One of the best experts on this subject based on the ideXlab platform.

  • the transformative power of the Thumbnail Image media logistics and infrastructural aesthetics
    First Monday, 2017
    Co-Authors: Nanna Bonde Thylstrup, Stina Teilmannlock
    Abstract:

    Thumbnail Images are discreet, yet central navigational tools in increasingly complex visual information environments. Indeed, without Thumbnail Images there would be no Image search: they are an inherent part of the information architecture of most digital information platforms. Yet, how might we understand the role of the Thumbnail as an attention technology in the digital economy? And what kind of aesthetic does it produce? This paper examines the legal negotiations of the Thumbnail Image and the ensuing decision to conceptualize the Thumbnail as a functional Image against the cultural history of visual attention technologies and the aesthetics of their connective function. Such an endeavour, we propose, allows us to understand and appreciate the significant digital economy and particular aesthetic of the Thumbnail Image despite its apparent subtlety.

Sue Thornton - One of the best experts on this subject based on the ideXlab platform.

  • video retrieval with the informedia digital video library system
    Text REtrieval Conference, 2002
    Co-Authors: Alexander G. Hauptmann, Rong Jin, Norman Papernick, Ricky Houghton, Sue Thornton
    Abstract:

    Background: The Informedia Digital Video Library System. The Informedia Digital Video Library [1] was the only NSF DLI project focusing specifically on information extraction from video and audio content. Over a terabyte of online data was collected, with automatically generated metadata and indices for retrieving videos from this library. The architecture for the project was based on the premise that real-time constraints on library and associated metadata creation could be relaxed in order to realize increased automation and deeper parsing and indexing for identifying the library contents and breaking it into segments. Library creation was an offline activity, with library exploration by users occurring online and making use of the generated metadata and segmentation. The goal of the Informedia interface was to enable quick access to relevant information in a digital video library, leveraging from derived metadata and the partitioning of the video into small segments. Figure 1 shows the IDVLS interface following a query. In this figure, a set of results is displayed at the bottom. The display includes a window containing a headline, and a pictorial menu of video segments each represented with a Thumbnail Image at approximately 1⁄4 resolution of the video in the horizontal and vertical dimensions. The headline window automatically pops up whenever the mouse is positioned over a result item; the headline window for the first result is shown. IDVLS also supports other ways of navigating and browsing the digital video library. These interface features were essential to deal with the ambiguity of the derived data generated by speech recognition, Image processing, and natural language processing. Consider the filmstrip and video playback IDVLS window shown in Figure 2. For this actual video in the IDVLS library, the segmentation process failed, resulting in a thirty-minute segment. This long segment was one of the returned results for the query “Mir collision.” The filmstrip in Figure 2 shows that the segment is more than just a story on the Russian space station, but rather begins with a commercial, then the weather, and then coverage of Hong Kong before addressing Mir. By overlaying the filmstrip and video playback windows with match location information, the user can quickly see that matches don’t occur until later in the segment, after these other stories that were irrelevant to the query. The match bars are optionally color-coded to specific query words; in Figure 2 “Mir” matches are in red and “collision” matches in purple. When the user moved the mouse over the match bars in the filmstrip, a text window displayed the actual matching word from the transcript or Video OCR metadata for that particular match; “Mir” is shown in one such text window in Figure 2. By investigating the distribution of match locations on the filmstrip, the user can determine the relevance of the returned result and the location of interest within the segment. The user can click on a match bar to jump directly to that point in the video segment. Hence, clicking the mouse as shown in Figure 2 would start playing the video at this mention of “Mir” with the overhead shot of people at desks. Similarly, IDVLS provided “seek to next match” and “seek to previous match” buttons in the video player allowing the user to quickly jump from one match to the next. In the example of Figure 2, these interface features allowed the user to bypass problems in segmentation and jump directly to the “Mir” story without having to first watch the opening video on other topics.

Nanna Bonde Thylstrup - One of the best experts on this subject based on the ideXlab platform.

  • the transformative power of the Thumbnail Image media logistics and infrastructural aesthetics
    First Monday, 2017
    Co-Authors: Nanna Bonde Thylstrup, Stina Teilmannlock
    Abstract:

    Thumbnail Images are discreet, yet central navigational tools in increasingly complex visual information environments. Indeed, without Thumbnail Images there would be no Image search: they are an inherent part of the information architecture of most digital information platforms. Yet, how might we understand the role of the Thumbnail as an attention technology in the digital economy? And what kind of aesthetic does it produce? This paper examines the legal negotiations of the Thumbnail Image and the ensuing decision to conceptualize the Thumbnail as a functional Image against the cultural history of visual attention technologies and the aesthetics of their connective function. Such an endeavour, we propose, allows us to understand and appreciate the significant digital economy and particular aesthetic of the Thumbnail Image despite its apparent subtlety.

Hungkhoon Tan - One of the best experts on this subject based on the ideXlab platform.

  • real time near duplicate elimination for web video search with content and context
    IEEE Transactions on Multimedia, 2009
    Co-Authors: Chongwah Ngo, Alexander G. Hauptmann, Hungkhoon Tan
    Abstract:

    With the exponential growth of social media, there exist huge numbers of near-duplicate web videos, ranging from simple formatting to complex mixture of different editing effects. In addition to the abundant video content, the social Web provides rich sets of context information associated with web videos, such as Thumbnail Image, time duration and so on. At the same time, the popularity of Web 2.0 demands for timely response to user queries. To balance the speed and accuracy aspects, in this paper, we combine the contextual information from time duration, number of views, and Thumbnail Images with the content analysis derived from color and local points to achieve real-time near-duplicate elimination. The results of 24 popular queries retrieved from YouTube show that the proposed approach integrating content and context can reach real-time novelty re-ranking of web videos with extremely high efficiency, where the majority of duplicates can be rapidly detected and removed from the top rankings. The speedup of the proposed approach can reach 164 times faster than the effective hierarchical method proposed in , with just a slight loss of performance.