Der Kern der Vorlesung 'Computer Vision for Human-Computer Interaction' ist das Finden und Verfolgen von Personen / Gesichtern in Bildern und Bildfolgen. Dabei werden folgende Themenfelder besprochen:
- Trackingverfahren: Kalman-Filter und Partikelfilter
Behandelter Stoff
| # | Datum | Kapitel | Inhalt |
|---|---|---|---|
| 1 | 18.10.2016 | Einführung | Organisatorisches und Überblick über den Stoff |
| 2 | 21.10.2016 | Klassifikation | Gaussian Mixture Models, EM, SVMs, Perceptron |
| 3 | 24.10.2016 | Face Detection I | Color Spaces (HSV, YUV), Histogram Backprojection, Histogram Matching, Mixture of Gaussians, ROC, Morphological Operations |
| 4 | 28.10.2016 | Face Detection II | Perceptron, MLP, Histogram Equalization, Haar-like features, Adaboost (Viola and Jones) |
| 5 | 31.10.2016 | Face Recognition I | Eigenface, Fisherface |
| 6 | 04.11.2016 | Programmierprojekte | Organisatorisches / Einführung dazu |
| 7 | 07.11.2016 | Face Recognition 2 | Alignment (range affine warps), Morphing models of 3d forms (PCA), |
| 8 | 11.11.2016 | CNNs | Convolution, Pooling, ReLU, Normalization layers, |
| 9 | 14.11.2016 | Programmierprojekte | Besprechung der praktischen Aufgaben |
| 10 | 18.11.2016 | ? | ? |
| 11 | 21.11.2016 | Facial Feature Detection | ? |
| 12 | 25.11.2016 | Automatic Facial Expression Analysis | ? |
| 13 | 28.11.2016 | Head Pose Estimation | Model-based approaches, Appearance-based approaches |
| 14 | 02.12.2016 | Person Detection | Introduction, HOG people detector, Silhouette matching |
| 15 | 05.12.2016 | ? | ? |
| 16 | 09.12.2016 | ? | ? |
| 17 | 16.12.2016 | Tracking II | ? |
| - | 20.01.2017 | Visuelle Perzeption | Kinect, Block Matching |
| - | 23.01.2017 | Gesture Recognition | HMMs; Pro Geste ein HMM trainieren |
| 18 | 27.01.2017 | Action & Activity Recognition I | ? |
| 19 | 30.01.2017 | Action & Activity Recognition II | ? |
| 20 | 06.02.2017 | Wrap-up | Zusammenfassung der wichtigsten Themen |
Klassifikatoren
- Classification
- SVM
- EM-Algorithmus (Expectation Maximization)
- Perceptron-Algorithmus
- k nearest neighbor (Mahalanobis-Distanz)
- Clustering
- Curse of Dimensionality
- Dimensionality reduction
- PCA
- LDA
- LDA (Linear discriminant analysis)
- Maximizes class separability (LDA)
Face Detection
- Face Detection
-
Face Detection ist die Aufgabe, in einem gegebenen Bild eine
Bounding-Box um jedes Gesicht zu zeichnen.
Ein Ansatz ist, zu versuchen, "Gesichtsfarbe" zu erkennen.
Die Modellierung kann mit Histogrammen erfolgen.
Datensätze:- ECU face detection database
- ECU face skin detection database
- MIT and CMU frontal face database
- Chromatic Color Spaces
- Nur zweidimensional (HS von HSV, UV von YUV, normalized rg von RGB). Soll robuster für die Erkennung von Hautfarbe sein.
- Histogram Backprojection
- Man geht für die Trainingsbeispiele alle Hautfarbe-Pixel durch. Bekommt man nun ein neues Bild, so geht man für dieses Bild jeden Pixel durch. Die Ausgabe ist ein Graustufenbild selber Größe, wo die Pixelfarbe die Anzahl der hautfarbenen Pixel dieser Farbe ist.
- Histogram Matching
-
Erstelle ein Histogramm für Hautfarbe. Um ein neues Bild zu
klassifizieren, bildet man Ausschnitte für das Bild. Für jeden
Ausschnitt vergleicht man das Histogramm des Ausschnitts mit dem
Hautfarbe-Histogramm. Dazu können folgende Metriken verwendet werden:
- Bhattacharyya distance
- Histogram intersection
- Earth-movers distance
- Histogram equalization
-
This method usually increases the global contrast of many images,
especially when the usable data of the image is represented by close
contrast values. Through this adjustment, the intensities can be better
distributed on the histogram. This allows for areas of lower local
contrast to gain a higher contrast. Histogram equalization accomplishes
this by effectively spreading out the most frequent intensity values.
(Source: Wikipedia)
The intensity values of the image are modified in such a way that the histogram is flattened.
It roughly works like this:- Let $p(x_i) = \frac{n_i}{n}$ be the probability of a pixel having level $i$.
- Let $c(i) = \sum_{j=0}^i p(x_j)$ be the cumulative distribution.
- Find transformation $T$ such that the cumulative distribution $y = T(x)$ is linear.
- Image normalization
-
Normalization is a process that changes the range of pixel intensity
values. Applications include photographs with poor contrast due to
glare, for example. Normalization is sometimes called contrast
stretching or histogram stretching.
(Source: Wikipedia) - ROC
- Die ROC-Kurve misst die Abwägung zwischen True-Positive-Rate ($\frac{TP}{Pos}$, y-Achse) und False Positive Rate ($\frac{FP}{Neg}$, x-Achse).
- Intersection over Union (IoU)
- Siehe StackExchange
- Haar-like features
- Based on Haar wavelets as features. Popularized by Viola and Jones (originally proposed by Papageorgiou et al.). A Haar-like feature considers adjacent rectangular regions at a specific location in a detection window, sums up the pixel intensities in each region and calculates the difference between these sums. Can be computed efficiently by using integral images.
- AdaBoost
- Siehe ML 1.
Face Recognition
Alignment:
- Works only with "nice" images (no occlusion, high resolution)
- Eyes are hard
- Computationally intensive (not suitable for real time applications)
- Face Recognition
- Face recognition has 4 main tasks:
- Face detection: Given an image, draw a rectangle around every face
- Face alignment: Transform a face to be in a canonical pose
- Face representation: Find a representation of a face which is suitable for follow-up tasks (small size, computationally cheap to compare, invariant to irrelevant changes)
- Face verification: Images of two faces are given. Decide if it is the same person or not.
Challenges:- Extrinsic Variations
- Illumination
- View-point
- Occlusion
- Imaging process (low resolution)
- Intrinsic Variations
- Aging
- Facial expressions
- Feature-Based (Geometric)
- fiducial points
- distances, angles, areas
- Appearance-Based
- holistic, fiducial regions
- statistical
- Face Recognition Tasks
-
- Closed set recognition: 120 Celebrities - who is it?
- Known / unknown: Is it one of the persons in the known set?
- Verification: Is it George Clooney?
- Open Set recognition: First known/unknown, then closed set.
- Gabor Filter
- A linear filter formed by a sinusoidal plane wave multiplied by a Gaussian envelope. It responds strongly to edges/textures of a specific orientation and frequency, similar to simple cells in the human visual cortex.
- Local Binary Pattern (LBP)
- For each pixel, compare its intensity to its neighbors (e.g. 8 neighbors in a 3×3 window): 1 if the neighbor is greater or equal, 0 otherwise. Reading these bits gives a code per pixel; a histogram of the codes over an image region is used as a texture descriptor.
- SIFT
- Detects and describes local keypoints that are invariant to scale, rotation and (partially) illumination. Keypoints are found via a Difference-of-Gaussians pyramid; each is described by a histogram of local gradient orientations. SIFT vector is 128 dimensional
- Bag of Visual Words (BoVW)
- Visual analogue of bag-of-words in NLP: local descriptors (e.g. SIFT) are clustered into a "visual vocabulary" (codebook); an image is represented as a histogram of how often each visual word occurs, discarding spatial information.
- Fisher Encoding
- Instead of plain BoVW counts, encodes the difference between local descriptors and a Gaussian Mixture Model fitted to the codebook (first- and second-order statistics). This gives a higher-dimensional, more discriminative representation than BoVW for the same vocabulary size.
CNNs
- Lernen Gabor Wavelets: die ersten Conv-Layer trainierter CNNs zeigen Gabor-artige und Farbblob-Filter (z.B. sichtbar in den AlexNet-Filtervisualisierungen von Krizhevsky et al., 2012)
- Warum tiefere Netze? → Muster können geteilt werden (Räder für Traktoren / Motorräder)
- Filter Response = Feature Map: die Ausgabe eines einzelnen Filters/Kernels auf den Input ist eine Feature Map
- Der Gradient von ReLU(x) für \(x \leq 0\) ist 0, aber das ist ok, weil pro Batch immer genug andere Units aktiv bleiben, über die weiter gelernt wird; anders als bei sigmoid/tanh verschwindet der Gradient nicht überall (Sättigung nur auf einer Seite)
- Negative Log Likelihood vs. Cross-Entropy: für eine Softmax-Ausgabe mit One-Hot-Ziel ist die Categorical-Cross-Entropy-Loss identisch zur NLL-Loss der wahren Klasse; Cross-Entropy vergleicht allgemein zwei Verteilungen, NLL ist die negative Log-Wahrscheinlichkeit der korrekten Klasse
- Pooling (Subsampling)
-
Auf eine $k \times k$ Region der Feature-Map wird ein Operator (z.B. max, mean)
angewendet, der diese zusammenfasst. Typischerweise wird diese Operation
mit einem Stride > 1 verwendet. Durch den Stride wird zugleich die
Datenmenge auf $\frac{1}{s^2}$ reduziert.
Pooling wird separat pro Feature-Map angewendet. - Normalization Layers
-
Normalize activations to stabilize training and improve generalization, either
locally within a feature map or by rescaling the value range. Batch Normalization
is a related, more commonly used technique that normalizes over the mini-batch.
- Contrast Normalization
- Range Normalization
Head-pose Estimation
- State of the Art: Regression based approach (Random Regression Forests) is 15 years old
- slide 14: \(r\) is distance, \(f\) is focal length, \(b\) is baseline, \(d_L, d_R\) is distance left/right - in the later slides is a figure.
- Pan / Tilt / Roll (bzw. Yaw / Pitch / Roll)
- Entropy
- The entropy $H$ of a discrete random variable $X$ is $$H(X) = \mathbb{E}[-\log_2(P(X))] = - \sum_{i=1}^n P(x_i) \log_2 P(x_i)$$
- Information Gain (Kullback–Leibler divergence)
- The expected information gain is the change in entropy $H$ from a prior state to a state that takes some information as given: $$IG(T,a) = H(T) - H(T|a)$$
- Disparitätenkarte
- Gibt für jeden Pixel eines Stereobildpaars den horizontalen Versatz (Disparität) zwischen linkem und rechtem Bild an; sie ist umgekehrt proportional zur Tiefe (näher = größere Disparität).
Facial Feature Detection
- Facial Features
-
- Nose tip
- Ears
- Eyes
- Chin
- Lips (corners)
- Eyebrow
- Facial Feature Detection
- Is important for alignment
Problems:
- Expression variations
- Scale variations / angle
- Lighting
- Occlusions
Statistical Appearance Models
- Represent shape and texture (learned separately): \(x \approx \bar{x} + P_s b_s\), where \(P_s\) are the eigenvectors of the covariance matrix; \(b_s = P_s^T (x-\bar{x})\)
Fit shape Model
- Active shape model (fitting algorithm)
Automatic Facial Expression Analysis
- Allgemein
-
- Valence-Arousal Modell von Russell 1980
- Facial Action Coding System (FACS)
- FACS (Facial Action Coding System)
- 44 Action Units, welche menschliche Mimik beschreiben
- 30 davon sind anatomisch bestimmten Gesichtsmuskeln zugeordnet (12 in der oberen, 18 in der unteren Gesichtshälfte)
- Viele binär, manche mit Intensität
- Emotionen sind Kombinationen daraus
- 44 Action Units, welche menschliche Mimik beschreiben
- Datenbanken
-
- Cohn-Kanade AU-Coded Facial Expression Database
- Emotion Recognition in the Wild (EmotiW)
- AFEW db
- Bag of Visual Words
- Siehe Bag of Visual Words oben.
Person detection
- Detection
- Detection is classification with localization
- Person detection
- Person detection (sometimes also pedestrian detection) is the
task of finding people in an image or video. This is usually done with
bounding boxes, although tight segmentation methods exist.
One can divide the methods the following way:- Input: Single image / video
- Detection approach
- Global: Detection of the person as a whole
- Part-based: Detection of arms, legs, head, ...
- Model type
- Generative: Models how the data was generated (+ is interpretable - hard to create)
- Discriminative
- L2 hysterese
- A variant of the L2 norm which cuts peaks
- HOG People Detector
- A person/pedestrian detector based on Histogram of Oriented Gradients features combined with a linear SVM classifier (Dalal & Triggs, 2005); a sliding-window, global, discriminative, single-image approach.
- Silhouette Matching
- Silhouette matching is practically not used anymore. It uses chamfer matching, 2/3 distance. It can be sped up with a template hierarchy.
Tracking
- Multi Camera Systems
- Topologies
- Stereo cameras (cars)
- wide baseline multi-camera system (conference room)
- non-overlapping fields of view (security system)
- TDOA (Time Difference of Arrival)
- In the case of multiple microphones and one audio source, the time delay when microphone 1 records the same as microphone 2 is called TDOA. It can be used to locate the audio source.
- Adaptive Merkmalsgewichtung
- Beim Tracking wird die Gewichtung einzelner Merkmale (z.B. Farbe, Form, Bewegung) laufend an die aktuelle Szene angepasst, um robuster gegen Beleuchtungsänderungen, Verdeckungen etc. zu sein.
Gesten
- Gesten
- Gesten bedeuten nicht überall dasselbe. Ein Nicken bedeutet in vielen, aber nicht in allen Kulturen Zustimmung (Quelle)
Action / Activity Recognition
- Aktion
- Zielgerichtete Interaktion mit der Umgebung; tendenziell mit wenigen oder Einzelschritten / kurzen Perioden
- Aktivität
- Sind aus vielen Einzelaktionen aufgebaut. Zwei verschiedene Aktivitäten
können aus sehr ähnlichen Bewegungen bestehen, aber deutlich
unterschiedliche Semantik haben. Ein Beispiel ist das Öffnen eines
Schlosses mit einem Schlüssel und das Herausdrehen einer Schraube
mit einem Schraubenzieher.
Es ist ein Klassifikationsproblem mit Bild-Sequenzen.
Lösungsansätze: HMMs, RNNs
Features: Histogramme des optischen Flusses- Global
- Lokal (Bild aufteilen)
- Implicit Shape Models
- Deskriptoren (SIFT)
Datasets: UCF Sport; Hollywood 2; Sports-1M - Zero-Crossing Rate
- Extrem einfaches Audio-Merkmal, welches es erlaubt, Auto-Hintergrundrauschen von Sprache zu unterscheiden. Die Zero-Crossing-Rate zählt, wie oft das Signal pro Zeiteinheit das Vorzeichen wechselt; Sprache und tonales Rauschen haben meist niedrigere ZCR-Werte als raues Hintergrundrauschen.
- MFCC (Mel Frequency Cepstral Coefficients)
- MFCCs sind gängige Features der Sprachverarbeitung.
- BoW (Bag of Words, Bag of Visual Words)
- Cluster Features for objects by feature similarity; gives histogram representation of object.
- HOG (Histogram of Oriented Gradients)
- Ein Merkmal für Bilder: Histogramme der Gradientenrichtungen in kleinen Zellen. SIFT-Deskriptoren verwenden eine ähnliche Idee.
- Optical Flow (OF)
- Das Vektorfeld, das für jeden Pixel angibt, wie er sich zwischen zwei aufeinanderfolgenden Bildern (scheinbar) bewegt hat, geschätzt aus der Änderung der Helligkeitsmuster.
- HOF (Histogram of Optical Flow)
- Histogramm der Richtungen (und ggf. Beträge) des optischen Flusses in einer lokalen Region; analog zu HOG, aber auf Bewegungsvektoren statt Gradienten angewendet.
- MBH (Motion Boundary Histogram)
- Wendet den HOG-Deskriptor auf die Ableitung des optischen Flusses statt auf den Fluss selbst an; dadurch werden konstante Kamerabewegungen herausgerechnet und nur relative Bewegung (Bewegungsränder) erfasst.
- Dense Trajectories
- Merkmal für Aktionserkennung in Videos: Punkte werden dicht in jedem Frame gesampelt und über mehrere Frames mittels optischem Fluss verfolgt (typischerweise 15 Frames); entlang der resultierenden Trajektorien werden HOG-, HOF- und MBH-Deskriptoren berechnet.
Wrap-up
- pinhole model, calibration (extrinsische / intrinsische Parameter), stereo processing (Disparitäten)
- Features
- Color, fg/bg, stereo edges, edge histogram, Gabor- und Haar-Filter (Viola&Jones), LBP
- Mid-level-Representations: GMM, BoW, BoW+Spatial layout (body parts)
- Dimensionalitätsreduktion: PCA (Eigenfaces) / LDA (Fisher-Faces)
- Classifiers: Boosting, SVM, CNN, Regression Trees, HMMs, k-NN
- Wie finde ich keypoints?
- Descriptors (SIFT!)
- Space-Time-Features: Space-Time-Interest points; HOG / HOF, Dense Trajectories: Bildfolge, finde Trajektorie
- Implicit Shape Model: Bag-of-Words + Spatial Layout
- Statistische Modelle: Active Shape / Active appearance
- Chromatische Farbräume: Helligkeit rausnormalisieren (2-dimensional: rg, HS, UV)
- Morphable 3D models: PCA - laser scans
- Active Shape: Statistische Modellierung der Form; Textur entlang der shape. Nutzt PCA. Lernt Textur.
- Active Appearance: 2D. Gemeinsames statistisches Modell von Shape & Appearance; Textur innerhalb von Mesh / Punkten. Nutzt auch PCA. Lernt Textur.
- 6 Basic Emotions: FACS, Action Units
- HOG modelliert Textur
- HOF modelliert Bewegung
Praktische Aufgaben
Es gibt 3 praktische Aufgaben, die 10% der Note ausmachen:
- Haut erkennen
- Detektieren, ob auf einem Bild eine Person ist oder nicht
- Erkennen, ob auf zwei gegebenen Bildern dieselbe Person ist
Die Aufgaben müssen mit C++ gemacht werden. OpenCV kann verwendet werden. Es ist Beispielcode gegeben.
Man muss in 180s die Modelle trainieren.
Bewertet wird eine Präsentation, die am 16.01.2016 gemacht werden muss. Die Präsentation soll mindestens 3 Folien, maximal 5 Folien haben. In diesen 5 Folien sollen alle 3 Aufgaben beschrieben werden. Es soll beschrieben werden, wie die Aufgaben gelöst wurden / was geklappt bzw. nicht geklappt hat. Die Präsentation soll ca. 8 - 10 min pro Team dauern.
Es ist ok, private Datensätze zu verwenden, um ggf. Hyperparameter zu bestimmen.
Für weitere Fragen steht Manuel Martinez zur Verfügung.
Prüfungsfragen
Was ist das Hauptproblem der Gesichtserkennung?
Welche farbbasierten Ansätze gibt es zur Gesichtserkennung?
Welche featurebasierten Ansätze gibt es zur Gesichtserkennung?
Wofür steht DCT und AAM?
PCA vs LDA: Was ist zur Face Recognition besser?
Wozu ist die Histogram Equalization gut?
Welche featurebasierten Ansätze gibt es zur Face Recognition?
Warum macht man Tracking und nicht einfach Frame-weise Detection?
Was sind Anwendungen von Action und Activity Recognition?
Was ist Kalibrierung?
Woher kommt die Skalen-Invarianz bei SIFT?
Woher kommt die Rotationsinvarianz bei SIFT?
Was ist der Unterschied zwischen diskriminativen und generativen Modellen?
Wie funktioniert Histogram Backprojection?
Welche Metriken kann man zum Benchmarken benutzen?
Was macht Computer Vision schwer?
Welche Annahmen macht man im Kalman-Filter, die man nicht im Particle-Filter hat?
Welche Anwendungen gibt es für CV in der HCI?
Auf welcher Ebene, wann und wie werden Informationen zusammengeführt (z.B. Face / Pose)?
Material und Links
- Vorlesungswebsite
- lecture-demo.ira.uka.de: Interaktive Demos, insbesondere Rosenblatt-Perceptron, ...
- StackExchange:
- Video Analysis lecture
Vorlesungsempfehlungen
Folgende Vorlesungen sind ähnlich:
- Analysetechniken großer Datenbestände
- Informationsfusion
- Machine Learning 1
- Machine Learning 2
- Mustererkennung
- Neuronale Netze
- Lokalisierung Mobiler Agenten
- Probabilistische Planung
Weitere:
- Einführung in die Bildfolgenauswertung
- Content-based Image and Video Retrieval - 3 ECTS, 2 SWS
Termine und Klausurablauf
Die Veranstaltung wird mündlich geprüft, jedoch sind 10% der Note durch praktische Aufgaben zu erlangen. Üblicherweise dauert eine Prüfung etwa 30 min.
Die Anmeldung zur Prüfung erfolgt per Email an das Sekretariat ([email protected]).
Weitere Prüfungstermine erst nach dem 18. April.