Audio Feature Extraction and Analysis for Scene Segmentation and Classification

Zhu Liu, Yao Wang, Tsuhan Chen

Research output: Contribution to journalArticlepeer-review


Understanding of the scene content of a video sequence is very important for content-based indexing and retrieval of multimedia databases. Research in this area in the past several years has focused on the use of speech recognition and image analysis techniques. As a complimentary effort to the prior work, we have focused on using the associated audio information (mainly the nonspeech portion) for video scene analysis. As an example, we consider the problem of discriminating five types of TV programs, namely commercials, basketball games, football games, news reports, and weather forecasts. A set of low-level audio features are proposed for characterizing semantic contents of short audio clips. The linear separability of different classes under the proposed feature space is examined using a clustering analysis. The effective features are identified by evaluating the intracluster and intercluster scattering matrices of the feature space. Using these features, a neural net classifier was successful in separating the above five types of TV programs. By evaluating the changes between the feature vectors of adjacent clips, we also can identify scene breaks in an audio sequence quite accurately. These results demonstrate the capability of the proposed audio features for characterizing the semantic content of an audio sequence.

Original languageEnglish (US)
Pages (from-to)61-79
Number of pages19
JournalJournal of VLSI Signal Processing Systems for Signal, Image, and Video Technology
Issue number1-2
StatePublished - 1998

ASJC Scopus subject areas

  • Signal Processing
  • Information Systems
  • Electrical and Electronic Engineering


Dive into the research topics of 'Audio Feature Extraction and Analysis for Scene Segmentation and Classification'. Together they form a unique fingerprint.

Cite this