Spaces:
Sleeping
Sleeping
Evan Li commited on
Commit Β·
8f19f34
1
Parent(s): ee3a08a
chopped
Browse files- Dockerfile +16 -4
- README.md +24 -17
- analyzers/aesthetic_analyzer.py +220 -0
- analyzers/beauty_analyzer.py +172 -0
- analyzers/demographic_analyzer.py +0 -227
- analyzers/ethnicity_analyzer.py +115 -0
- analyzers/insightface_analyzer.py +171 -0
- analyzers/landmark_analyzer.py +3 -0
- analyzers/parsing_analyzer.py +6 -1
- app.py +215 -154
- architecture.md +101 -56
- requirements.txt +2 -0
Dockerfile
CHANGED
|
@@ -2,7 +2,7 @@ FROM python:3.11-slim
|
|
| 2 |
|
| 3 |
# System deps for OpenCV and MediaPipe
|
| 4 |
RUN apt-get update && apt-get install -y --no-install-recommends \
|
| 5 |
-
libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 wget \
|
| 6 |
libegl1 libgles2 libgomp1 \
|
| 7 |
g++ && \
|
| 8 |
rm -rf /var/lib/apt/lists/*
|
|
@@ -14,13 +14,25 @@ COPY requirements.txt .
|
|
| 14 |
RUN pip install --no-cache-dir -r requirements.txt
|
| 15 |
|
| 16 |
# Pre-download MediaPipe model at build time so first request is fast.
|
| 17 |
-
#
|
| 18 |
-
# HairTypeViT)
|
| 19 |
-
# in /root/.cache
|
|
|
|
| 20 |
RUN mkdir -p models && \
|
| 21 |
wget -q -O models/face_landmarker.task \
|
| 22 |
"https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/latest/face_landmarker.task"
|
| 23 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
COPY . .
|
| 25 |
|
| 26 |
EXPOSE 7860
|
|
|
|
| 2 |
|
| 3 |
# System deps for OpenCV and MediaPipe
|
| 4 |
RUN apt-get update && apt-get install -y --no-install-recommends \
|
| 5 |
+
libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 wget unzip \
|
| 6 |
libegl1 libgles2 libgomp1 \
|
| 7 |
g++ && \
|
| 8 |
rm -rf /var/lib/apt/lists/*
|
|
|
|
| 14 |
RUN pip install --no-cache-dir -r requirements.txt
|
| 15 |
|
| 16 |
# Pre-download MediaPipe model at build time so first request is fast.
|
| 17 |
+
# Hugging Face models (Ethnicity ViT, SegFormer, HSEmotion, ObstructionViT,
|
| 18 |
+
# HairTypeViT) and the InsightFace buffalo_l bundle are pulled lazily on
|
| 19 |
+
# first request and cached in /root/.cache for the lifetime of the
|
| 20 |
+
# container.
|
| 21 |
RUN mkdir -p models && \
|
| 22 |
wget -q -O models/face_landmarker.task \
|
| 23 |
"https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/latest/face_landmarker.task"
|
| 24 |
|
| 25 |
+
# Pre-download InsightFace buffalo_l bundle (detection + recognition +
|
| 26 |
+
# age + gender + landmarks) so the first /analyze call doesn't pay the
|
| 27 |
+
# ~280MB download. The bundle auto-extracts under ~/.insightface/models/
|
| 28 |
+
# on first use.
|
| 29 |
+
RUN mkdir -p /root/.insightface/models && \
|
| 30 |
+
wget -q -O /root/.insightface/models/buffalo_l.zip \
|
| 31 |
+
"https://github.com/deepinsight/insightface/releases/download/v0.7/buffalo_l.zip" && \
|
| 32 |
+
cd /root/.insightface/models && unzip -q buffalo_l.zip -d buffalo_l && rm buffalo_l.zip
|
| 33 |
+
|
| 34 |
+
# unzip wasn't in the system deps; add it via the apt block at the top.
|
| 35 |
+
|
| 36 |
COPY . .
|
| 37 |
|
| 38 |
EXPOSE 7860
|
README.md
CHANGED
|
@@ -10,26 +10,33 @@ pinned: false
|
|
| 10 |
|
| 11 |
# HCP Face Analysis Microservice
|
| 12 |
|
| 13 |
-
FastAPI service that runs
|
| 14 |
-
and returns a merged dictionary of
|
|
|
|
| 15 |
|
| 16 |
## Models
|
| 17 |
|
| 18 |
| # | Component | Model | Task | Size |
|
| 19 |
|---|-----------|-------|------|------|
|
| 20 |
-
| 1 |
|
| 21 |
-
| 2 |
|
| 22 |
-
|
|
| 23 |
-
|
|
| 24 |
-
|
|
| 25 |
-
|
|
| 26 |
-
|
|
| 27 |
-
|
|
| 28 |
-
|
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
## API endpoints
|
| 35 |
|
|
@@ -46,5 +53,5 @@ curl -X POST https://YOUR-SPACE.hf.space/analyze-base64 \
|
|
| 46 |
-d '{"image": "<base64-encoded-image>"}'
|
| 47 |
```
|
| 48 |
|
| 49 |
-
See [architecture.md](./architecture.md) for the pipeline diagram and
|
| 50 |
-
full per-attribute
|
|
|
|
| 10 |
|
| 11 |
# HCP Face Analysis Microservice
|
| 12 |
|
| 13 |
+
FastAPI service that runs ten specialised analyzers over a single
|
| 14 |
+
photo and returns a merged dictionary of facial attributes plus a
|
| 15 |
+
face-recognition embedding and an aesthetic "chopped score."
|
| 16 |
|
| 17 |
## Models
|
| 18 |
|
| 19 |
| # | Component | Model | Task | Size |
|
| 20 |
|---|-----------|-------|------|------|
|
| 21 |
+
| 1 | InsightFace | `buffalo_l` (SCRFD + ArcFace ResNet50, ONNX) | Detection + 512-d recognition embedding + age + gender + 106 landmarks (99.83% LFW) | ~280 MB |
|
| 22 |
+
| 2 | MediaPipe Face Landmarker | `face_landmarker.task` (Google) | 478 3D landmarks + 52 ARKit blendshapes β geometric features, smiling, mouth-open | ~4 MB |
|
| 23 |
+
| 3 | Ethnicity | `cledoux42/Ethnicity_Test_v003` (ViT) | 5-class ethnicity (~79.6% acc) | ~340 MB |
|
| 24 |
+
| 4 | Human parsing | `matei-dorian/segformer-b5-finetuned-human-parsing` | 18-class pixel segmentation β masks + hair length + hat | ~340 MB |
|
| 25 |
+
| 5 | Emotion | HSEmotion `enet_b0_8_best_afew` (EfficientNet-B0) | 8-class emotion + valence/arousal | ~20 MB |
|
| 26 |
+
| 6 | Color analysis | (no model β OpenCV LAB/HSV) | Skin tone, hair color, eye color, lip color | 0 MB |
|
| 27 |
+
| 7 | Obstruction | `dima806/face_obstruction_image_detection` (ViT-B/16) | glasses / sunglasses / mask (~99% precision) | ~340 MB |
|
| 28 |
+
| 8 | Hair type | `dima806/hair_type_image_detection` (ViT-B/16) | curly/dreadlocks/kinky/straight/wavy (~93% acc) | ~340 MB |
|
| 29 |
+
| 9 | Beauty regression | timm ResNet-50, fine-tuned on SCUT-FBP5500 | 1.0β5.0 beauty score (~Pearson r β₯ 0.85 expected) | ~100 MB |
|
| 30 |
+
| 10 | Aesthetic aggregator | (no model β Python rules) | Combines learned beauty + heuristic factors β chopped_score (0β100) | 0 MB |
|
| 31 |
+
|
| 32 |
+
InsightFace `buffalo_l` is pre-downloaded at Docker build time. The
|
| 33 |
+
MediaPipe weight file is pre-downloaded too. All Hugging Face models
|
| 34 |
+
lazy-load on first inference and cache locally for the lifetime of
|
| 35 |
+
the process. The SCUT-FBP5500-trained beauty regressor must be
|
| 36 |
+
produced via [`training/beauty/`](../training/beauty/) β until you
|
| 37 |
+
drop weights in `models/beauty_regressor.pt` (or set
|
| 38 |
+
`BEAUTY_HF_REPO_ID`), `beauty_score` returns null and the chopped
|
| 39 |
+
score uses heuristics only.
|
| 40 |
|
| 41 |
## API endpoints
|
| 42 |
|
|
|
|
| 53 |
-d '{"image": "<base64-encoded-image>"}'
|
| 54 |
```
|
| 55 |
|
| 56 |
+
See [architecture.md](./architecture.md) for the pipeline diagram and
|
| 57 |
+
the full per-attribute β source map.
|
analyzers/aesthetic_analyzer.py
ADDED
|
@@ -0,0 +1,220 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
AestheticAnalyzer β "chopped score" aggregator.
|
| 3 |
+
|
| 4 |
+
What it does
|
| 5 |
+
------------
|
| 6 |
+
Reads the merged result dict from every other analyzer and produces a
|
| 7 |
+
single numeric "chopped score" plus a per-factor breakdown. Higher
|
| 8 |
+
score = more chopped = less conventionally attractive (by the
|
| 9 |
+
arbitrary rubric encoded here). The breakdown lets you tune weights
|
| 10 |
+
or flip polarity client-side without rerunning inference.
|
| 11 |
+
|
| 12 |
+
Score composition
|
| 13 |
+
-----------------
|
| 14 |
+
The final chopped_score is a weighted blend of two sources:
|
| 15 |
+
|
| 16 |
+
1. **Learned beauty regressor** (from BeautyAnalyzer, trained on
|
| 17 |
+
SCUT-FBP5500): a number in [1.0, 5.0] reflecting averaged human
|
| 18 |
+
ratings. We rescale to a 0β100 unattractiveness axis. This is the
|
| 19 |
+
dominant signal when available β heavy weight (default 0.7).
|
| 20 |
+
|
| 21 |
+
2. **Rule-based factor sum**: penalties for asymmetry, wrinkles,
|
| 22 |
+
uneven skin, freckles, and asymmetric smile; bonuses for defined
|
| 23 |
+
jawline, prominent cheekbones, clear skin, balanced lips, and
|
| 24 |
+
dimples. Each factor is documented in `_compute_rule_score`.
|
| 25 |
+
This is the only signal when the regressor isn't loaded
|
| 26 |
+
(BeautyAnalyzer returns None).
|
| 27 |
+
|
| 28 |
+
Blend math
|
| 29 |
+
----------
|
| 30 |
+
if beauty_score available:
|
| 31 |
+
chopped = 0.7 * (100 - beauty_norm) + 0.3 * rule_score
|
| 32 |
+
else:
|
| 33 |
+
chopped = rule_score
|
| 34 |
+
chopped is clamped to [0, 100].
|
| 35 |
+
|
| 36 |
+
Subjectivity disclaimer
|
| 37 |
+
-----------------------
|
| 38 |
+
Every weight in this file is a guess. "Beauty" is subjective, culturally
|
| 39 |
+
biased, and reductive. Treat the score as an in-joke metric; never
|
| 40 |
+
expose it as objective truth. The UI gates the row behind a
|
| 41 |
+
Settings toggle off-by-default for that reason.
|
| 42 |
+
|
| 43 |
+
Note: this analyzer takes no image input β it reads the merged result
|
| 44 |
+
dict produced by every other analyzer that ran ahead of it.
|
| 45 |
+
"""
|
| 46 |
+
|
| 47 |
+
from typing import Any
|
| 48 |
+
|
| 49 |
+
|
| 50 |
+
# How much weight the learned beauty regressor gets when both signals
|
| 51 |
+
# are available. The rule-based sum gets the rest (1 - this).
|
| 52 |
+
LEARNED_WEIGHT = 0.7
|
| 53 |
+
|
| 54 |
+
# Baseline score. Penalties push up, bonuses pull down.
|
| 55 |
+
BASELINE = 50.0
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
class AestheticAnalyzer:
|
| 59 |
+
def __init__(self):
|
| 60 |
+
# No model to load.
|
| 61 |
+
pass
|
| 62 |
+
|
| 63 |
+
def analyze(self, merged: dict[str, Any]) -> dict[str, Any]:
|
| 64 |
+
"""Compute chopped score from the merged result dict.
|
| 65 |
+
|
| 66 |
+
Unusual signature: not (img_rgb), since this analyzer aggregates
|
| 67 |
+
prior results rather than running inference on the image.
|
| 68 |
+
app.py special-cases the call to pass `merged` here.
|
| 69 |
+
"""
|
| 70 |
+
rule_score, breakdown = self._compute_rule_score(merged)
|
| 71 |
+
|
| 72 |
+
beauty_norm = merged.get("beauty_score_norm")
|
| 73 |
+
if beauty_norm is not None:
|
| 74 |
+
# Beauty regressor: 0 = ugly, 100 = beautiful (per SCUT-FBP5500
|
| 75 |
+
# scaling). Flip to unattractiveness axis: 100 - x.
|
| 76 |
+
learned_unattractive = 100.0 - float(beauty_norm)
|
| 77 |
+
chopped = (
|
| 78 |
+
LEARNED_WEIGHT * learned_unattractive
|
| 79 |
+
+ (1.0 - LEARNED_WEIGHT) * rule_score
|
| 80 |
+
)
|
| 81 |
+
breakdown["learned_unattractive"] = round(
|
| 82 |
+
LEARNED_WEIGHT * learned_unattractive - LEARNED_WEIGHT * BASELINE, 2
|
| 83 |
+
)
|
| 84 |
+
breakdown["_blend_weight_learned"] = LEARNED_WEIGHT
|
| 85 |
+
else:
|
| 86 |
+
chopped = rule_score
|
| 87 |
+
breakdown["_blend_weight_learned"] = 0.0
|
| 88 |
+
|
| 89 |
+
chopped = max(0.0, min(100.0, chopped))
|
| 90 |
+
|
| 91 |
+
return {
|
| 92 |
+
"chopped_score": round(chopped, 1),
|
| 93 |
+
"chopped_breakdown": breakdown,
|
| 94 |
+
"chopped_polarity_note": (
|
| 95 |
+
"0 = least chopped, 100 = most chopped. "
|
| 96 |
+
"Subtract from 100 for an 'attractiveness' read."
|
| 97 |
+
),
|
| 98 |
+
}
|
| 99 |
+
|
| 100 |
+
# ------------------------------------------------------------------
|
| 101 |
+
# Rule-based scoring
|
| 102 |
+
# ------------------------------------------------------------------
|
| 103 |
+
|
| 104 |
+
@staticmethod
|
| 105 |
+
def _compute_rule_score(d: dict[str, Any]) -> tuple[float, dict[str, float]]:
|
| 106 |
+
"""Hand-tuned weighted sum over previously-extracted attributes.
|
| 107 |
+
|
| 108 |
+
Returns (score, breakdown_dict). The breakdown gives each factor's
|
| 109 |
+
signed contribution so a UI can show *why* a score landed where
|
| 110 |
+
it did. Score starts at BASELINE (50) and moves up/down.
|
| 111 |
+
"""
|
| 112 |
+
score = BASELINE
|
| 113 |
+
breakdown: dict[str, float] = {}
|
| 114 |
+
|
| 115 |
+
# ββ Penalties (push score up = more chopped) βββββββββββββββββ
|
| 116 |
+
|
| 117 |
+
# Facial asymmetry: 0 = perfectly symmetric, 1 = very asymmetric.
|
| 118 |
+
# MediaPipe `facial_asymmetry_score` is already in this range.
|
| 119 |
+
asym = d.get("facial_asymmetry_score")
|
| 120 |
+
if isinstance(asym, (int, float)):
|
| 121 |
+
penalty = float(asym) * 18.0
|
| 122 |
+
score += penalty
|
| 123 |
+
breakdown["asymmetry_penalty"] = round(penalty, 2)
|
| 124 |
+
|
| 125 |
+
# Wrinkle level from SegFormer + OpenCV Laplacian classification.
|
| 126 |
+
wrinkle_penalty_map = {
|
| 127 |
+
"smooth": 0.0, "slight": 4.0, "moderate": 8.0, "prominent": 12.0,
|
| 128 |
+
}
|
| 129 |
+
wrinkle = d.get("wrinkle_level")
|
| 130 |
+
if wrinkle in wrinkle_penalty_map:
|
| 131 |
+
penalty = wrinkle_penalty_map[wrinkle]
|
| 132 |
+
score += penalty
|
| 133 |
+
breakdown["wrinkle_penalty"] = penalty
|
| 134 |
+
|
| 135 |
+
# Skin uniformity = LAB L* std-dev over the face mask. Higher
|
| 136 |
+
# std means uneven tone (shadows, blemishes). Scale up to +8.
|
| 137 |
+
uniformity = d.get("skin_uniformity")
|
| 138 |
+
if isinstance(uniformity, (int, float)) and uniformity > 0:
|
| 139 |
+
# Empirically, uniformity in clean skin is ~8-15; very uneven
|
| 140 |
+
# skin pushes into the 20-30 range.
|
| 141 |
+
penalty = min(8.0, max(0.0, (float(uniformity) - 10.0) * 0.5))
|
| 142 |
+
score += penalty
|
| 143 |
+
breakdown["skin_unevenness_penalty"] = round(penalty, 2)
|
| 144 |
+
|
| 145 |
+
# Freckles/moles bucket.
|
| 146 |
+
freckle_penalty_map = {"none": 0.0, "few": 1.0, "some": 3.0, "many": 5.0}
|
| 147 |
+
freckles = d.get("freckles_or_moles")
|
| 148 |
+
if freckles in freckle_penalty_map:
|
| 149 |
+
penalty = freckle_penalty_map[freckles]
|
| 150 |
+
score += penalty
|
| 151 |
+
breakdown["freckles_penalty"] = penalty
|
| 152 |
+
|
| 153 |
+
# Smile asymmetry: 0 = perfectly symmetric smile, larger = lopsided.
|
| 154 |
+
smile_asym = d.get("smile_asymmetry")
|
| 155 |
+
if isinstance(smile_asym, (int, float)):
|
| 156 |
+
penalty = min(6.0, float(smile_asym) * 30.0)
|
| 157 |
+
score += penalty
|
| 158 |
+
breakdown["smile_asymmetry_penalty"] = round(penalty, 2)
|
| 159 |
+
|
| 160 |
+
# Photo-quality penalty: sunglasses/mask hide features and the
|
| 161 |
+
# model is guessing more. Mild penalty, not a personal trait.
|
| 162 |
+
if d.get("wearing_sunglasses") or d.get("wearing_mask"):
|
| 163 |
+
score += 5.0
|
| 164 |
+
breakdown["obstruction_penalty"] = 5.0
|
| 165 |
+
|
| 166 |
+
# ββ Bonuses (pull score down = less chopped) βββββββββββββββββ
|
| 167 |
+
|
| 168 |
+
# Defined jawline. Two signals (string bucket + numeric angle);
|
| 169 |
+
# take the stronger of the two contributions.
|
| 170 |
+
jaw_bonus = 0.0
|
| 171 |
+
jaw_type = d.get("jawline_type")
|
| 172 |
+
jaw_type_bonus_map = {"sharp": -10.0, "strong": -6.0, "soft": 0.0}
|
| 173 |
+
if jaw_type in jaw_type_bonus_map:
|
| 174 |
+
jaw_bonus = jaw_type_bonus_map[jaw_type]
|
| 175 |
+
jaw_angle = d.get("jawline_angle")
|
| 176 |
+
if isinstance(jaw_angle, (int, float)) and jaw_angle < 115:
|
| 177 |
+
# Sharp angles add more on top of the categorical signal.
|
| 178 |
+
jaw_bonus = min(jaw_bonus, -10.0)
|
| 179 |
+
if jaw_bonus:
|
| 180 |
+
score += jaw_bonus
|
| 181 |
+
breakdown["jaw_definition_bonus"] = round(jaw_bonus, 2)
|
| 182 |
+
|
| 183 |
+
# Cheekbone prominence.
|
| 184 |
+
cheek_bonus_map = {"high": -7.0, "moderate": -3.0, "flat": 0.0}
|
| 185 |
+
cheek = d.get("cheekbone_prominence")
|
| 186 |
+
if cheek in cheek_bonus_map:
|
| 187 |
+
bonus = cheek_bonus_map[cheek]
|
| 188 |
+
score += bonus
|
| 189 |
+
breakdown["cheekbone_bonus"] = bonus
|
| 190 |
+
|
| 191 |
+
# Skin clarity bonus when the texture score is low (i.e. smooth skin).
|
| 192 |
+
# skin_texture_score is the same Laplacian-density value used by
|
| 193 |
+
# wrinkle_level; β€4 is "smooth" territory.
|
| 194 |
+
texture = d.get("skin_texture_score")
|
| 195 |
+
if isinstance(texture, (int, float)) and 0 < texture <= 4:
|
| 196 |
+
score -= 9.0
|
| 197 |
+
breakdown["skin_clarity_bonus"] = -9.0
|
| 198 |
+
|
| 199 |
+
# Lip fullness β "average" and "full" both read as healthy.
|
| 200 |
+
lip = d.get("lip_fullness")
|
| 201 |
+
if lip in {"average", "full"}:
|
| 202 |
+
score -= 5.0
|
| 203 |
+
breakdown["lip_fullness_bonus"] = -5.0
|
| 204 |
+
|
| 205 |
+
# Defined cupid's bow.
|
| 206 |
+
if d.get("cupids_bow") == "defined":
|
| 207 |
+
score -= 3.0
|
| 208 |
+
breakdown["cupids_bow_bonus"] = -3.0
|
| 209 |
+
|
| 210 |
+
# Normal eye spacing.
|
| 211 |
+
if d.get("eye_spacing") == "average":
|
| 212 |
+
score -= 4.0
|
| 213 |
+
breakdown["eye_spacing_bonus"] = -4.0
|
| 214 |
+
|
| 215 |
+
# Dimples β small bonus when the MediaPipe heuristic fires.
|
| 216 |
+
if d.get("possible_dimples"):
|
| 217 |
+
score -= 3.0
|
| 218 |
+
breakdown["dimples_bonus"] = -3.0
|
| 219 |
+
|
| 220 |
+
return score, breakdown
|
analyzers/beauty_analyzer.py
ADDED
|
@@ -0,0 +1,172 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
BeautyAnalyzer β learned facial-beauty regression on SCUT-FBP5500.
|
| 3 |
+
|
| 4 |
+
Model
|
| 5 |
+
-----
|
| 6 |
+
- Architecture : timm ResNet-50 with a single-output regression head
|
| 7 |
+
(output value in the SCUT-FBP5500 1.0β5.0 score range,
|
| 8 |
+
averaged from 60 human raters per image).
|
| 9 |
+
- Dataset : SCUT-FBP5500 (5,500 faces, gender + race balanced).
|
| 10 |
+
https://github.com/HCIILAB/SCUT-FBP5500-Database-Release
|
| 11 |
+
- Trained by : the user, via the training kit at `training/beauty/`.
|
| 12 |
+
- Expected MAE : ~0.25β0.30 on the standard test split; Pearson r β₯ 0.85.
|
| 13 |
+
- License : training data is research-only; the trained weights
|
| 14 |
+
you produce are yours to license as you choose.
|
| 15 |
+
|
| 16 |
+
Weight loading
|
| 17 |
+
--------------
|
| 18 |
+
Two ways the analyzer finds weights, tried in order:
|
| 19 |
+
1. Local file at `models/beauty_regressor.pt` (drop in after training).
|
| 20 |
+
2. Hugging Face Hub repo, controlled by `BEAUTY_HF_REPO_ID` env var
|
| 21 |
+
(e.g. `your-username/scut-fbp5500-resnet50`), loaded via
|
| 22 |
+
huggingface_hub.hf_hub_download.
|
| 23 |
+
|
| 24 |
+
If neither resolves, the analyzer logs a warning and returns
|
| 25 |
+
`beauty_score: None`, which AestheticAnalyzer detects and falls back
|
| 26 |
+
to the pure rule-based chopped score.
|
| 27 |
+
|
| 28 |
+
Inputs
|
| 29 |
+
------
|
| 30 |
+
img_rgb : np.ndarray (H, W, 3) uint8
|
| 31 |
+
|
| 32 |
+
Outputs (dict)
|
| 33 |
+
--------------
|
| 34 |
+
beauty_score : float in [1.0, 5.0] (SCUT-FBP5500 native range)
|
| 35 |
+
or None if no model is available
|
| 36 |
+
beauty_score_norm : float in [0.0, 100.0] (linearly rescaled)
|
| 37 |
+
beauty_model_source : "local" | "huggingface" | "unavailable"
|
| 38 |
+
"""
|
| 39 |
+
|
| 40 |
+
import os
|
| 41 |
+
from typing import Any
|
| 42 |
+
|
| 43 |
+
import numpy as np
|
| 44 |
+
from PIL import Image
|
| 45 |
+
|
| 46 |
+
try:
|
| 47 |
+
import torch
|
| 48 |
+
import timm
|
| 49 |
+
from torchvision import transforms
|
| 50 |
+
HAS_TORCH = True
|
| 51 |
+
except ImportError:
|
| 52 |
+
HAS_TORCH = False
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
LOCAL_WEIGHTS_PATH = os.environ.get(
|
| 56 |
+
"BEAUTY_WEIGHTS_PATH", "models/beauty_regressor.pt"
|
| 57 |
+
)
|
| 58 |
+
HF_REPO_ID = os.environ.get("BEAUTY_HF_REPO_ID") # e.g. "user/scut-fbp5500-resnet50"
|
| 59 |
+
HF_FILENAME = os.environ.get("BEAUTY_HF_FILENAME", "beauty_regressor.pt")
|
| 60 |
+
BACKBONE = os.environ.get("BEAUTY_BACKBONE", "resnet50")
|
| 61 |
+
|
| 62 |
+
# Standard ImageNet stats β SCUT-FBP5500 fine-tunes from ImageNet-pretrained
|
| 63 |
+
# backbones so we use the same normalisation at inference time.
|
| 64 |
+
IMAGENET_MEAN = [0.485, 0.456, 0.406]
|
| 65 |
+
IMAGENET_STD = [0.229, 0.224, 0.225]
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
class BeautyAnalyzer:
|
| 69 |
+
def __init__(self):
|
| 70 |
+
self.device = (
|
| 71 |
+
torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
| 72 |
+
if HAS_TORCH else None
|
| 73 |
+
)
|
| 74 |
+
self.model = None
|
| 75 |
+
self.source = "unavailable"
|
| 76 |
+
self.transform = None
|
| 77 |
+
|
| 78 |
+
if not HAS_TORCH:
|
| 79 |
+
print(
|
| 80 |
+
"[BeautyAnalyzer] torch / timm not installed β beauty_score "
|
| 81 |
+
"will be None and AestheticAnalyzer will fall back to rules."
|
| 82 |
+
)
|
| 83 |
+
return
|
| 84 |
+
|
| 85 |
+
weights_path = self._resolve_weights_path()
|
| 86 |
+
if weights_path is None:
|
| 87 |
+
print(
|
| 88 |
+
"[BeautyAnalyzer] No trained weights found at "
|
| 89 |
+
f"{LOCAL_WEIGHTS_PATH} and BEAUTY_HF_REPO_ID is unset. "
|
| 90 |
+
"Train one via `training/beauty/train.py` and drop the "
|
| 91 |
+
".pt into face-service/models/, or set BEAUTY_HF_REPO_ID."
|
| 92 |
+
)
|
| 93 |
+
return
|
| 94 |
+
|
| 95 |
+
try:
|
| 96 |
+
# Build backbone with a single regression output. Matches the
|
| 97 |
+
# training script's architecture exactly β see
|
| 98 |
+
# training/beauty/train.py.
|
| 99 |
+
self.model = timm.create_model(BACKBONE, pretrained=False, num_classes=1)
|
| 100 |
+
state = torch.load(weights_path, map_location=self.device)
|
| 101 |
+
# Support both bare state_dicts and {"state_dict": ...} wrappers.
|
| 102 |
+
if isinstance(state, dict) and "state_dict" in state:
|
| 103 |
+
state = state["state_dict"]
|
| 104 |
+
self.model.load_state_dict(state, strict=True)
|
| 105 |
+
self.model.to(self.device).eval()
|
| 106 |
+
|
| 107 |
+
# SCUT-FBP5500 standard inference transform: 224Γ224, ImageNet norm.
|
| 108 |
+
self.transform = transforms.Compose([
|
| 109 |
+
transforms.Resize((224, 224)),
|
| 110 |
+
transforms.ToTensor(),
|
| 111 |
+
transforms.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD),
|
| 112 |
+
])
|
| 113 |
+
print(
|
| 114 |
+
f"[BeautyAnalyzer] Loaded {BACKBONE} weights from "
|
| 115 |
+
f"{weights_path} (source={self.source})"
|
| 116 |
+
)
|
| 117 |
+
except Exception as exc:
|
| 118 |
+
print(f"[BeautyAnalyzer] Failed to load beauty model: {exc}")
|
| 119 |
+
self.model = None
|
| 120 |
+
|
| 121 |
+
@staticmethod
|
| 122 |
+
def _resolve_weights_path() -> str | None:
|
| 123 |
+
"""Local file wins, HF Hub is the fallback."""
|
| 124 |
+
if os.path.exists(LOCAL_WEIGHTS_PATH):
|
| 125 |
+
return LOCAL_WEIGHTS_PATH
|
| 126 |
+
if HF_REPO_ID:
|
| 127 |
+
try:
|
| 128 |
+
from huggingface_hub import hf_hub_download
|
| 129 |
+
return hf_hub_download(repo_id=HF_REPO_ID, filename=HF_FILENAME)
|
| 130 |
+
except Exception as exc:
|
| 131 |
+
print(f"[BeautyAnalyzer] HF Hub download failed: {exc}")
|
| 132 |
+
return None
|
| 133 |
+
|
| 134 |
+
def _set_source(self, source: str) -> None:
|
| 135 |
+
self.source = source
|
| 136 |
+
|
| 137 |
+
def analyze(self, img_rgb: np.ndarray) -> dict[str, Any]:
|
| 138 |
+
if self.model is None or self.transform is None:
|
| 139 |
+
return self._empty_result()
|
| 140 |
+
|
| 141 |
+
try:
|
| 142 |
+
pil = Image.fromarray(img_rgb).convert("RGB")
|
| 143 |
+
tensor = self.transform(pil).unsqueeze(0).to(self.device)
|
| 144 |
+
# Wrap inference in no_grad to save memory; HAS_TORCH is
|
| 145 |
+
# guaranteed True here because self.model wouldn't exist
|
| 146 |
+
# otherwise.
|
| 147 |
+
with torch.no_grad():
|
| 148 |
+
raw = self.model(tensor).squeeze().item() # scalar in ~[1, 5]
|
| 149 |
+
except Exception as exc:
|
| 150 |
+
print(f"[BeautyAnalyzer] Inference failed: {exc}")
|
| 151 |
+
return self._empty_result()
|
| 152 |
+
|
| 153 |
+
# Clamp to SCUT-FBP5500's nominal 1β5 range; out-of-range outputs
|
| 154 |
+
# mean the regressor's extrapolating beyond its training labels.
|
| 155 |
+
score = max(1.0, min(5.0, float(raw)))
|
| 156 |
+
|
| 157 |
+
# Linear rescale to 0β100 for downstream display. 1.0 β 0, 5.0 β 100.
|
| 158 |
+
norm = (score - 1.0) / 4.0 * 100.0
|
| 159 |
+
|
| 160 |
+
return {
|
| 161 |
+
"beauty_score": round(score, 3),
|
| 162 |
+
"beauty_score_norm": round(norm, 1),
|
| 163 |
+
"beauty_model_source": self.source or "local",
|
| 164 |
+
}
|
| 165 |
+
|
| 166 |
+
@staticmethod
|
| 167 |
+
def _empty_result() -> dict[str, Any]:
|
| 168 |
+
return {
|
| 169 |
+
"beauty_score": None,
|
| 170 |
+
"beauty_score_norm": None,
|
| 171 |
+
"beauty_model_source": "unavailable",
|
| 172 |
+
}
|
analyzers/demographic_analyzer.py
DELETED
|
@@ -1,227 +0,0 @@
|
|
| 1 |
-
"""
|
| 2 |
-
DemographicAnalyzer β age, gender, ethnicity via three ViT classifiers.
|
| 3 |
-
|
| 4 |
-
Models
|
| 5 |
-
------
|
| 6 |
-
- Age : dima806/fairface_age_image_detection
|
| 7 |
-
ViT-B/16, ~59% top-1 on FairFace 9 age buckets.
|
| 8 |
-
- Gender : dima806/fairface_gender_image_detection
|
| 9 |
-
ViT-B/16, ~93.4% on FairFace.
|
| 10 |
-
- Ethnicity : cledoux42/Ethnicity_Test_v003
|
| 11 |
-
ViT, 79.6% accuracy, macro-F1 0.797. 5-class output that
|
| 12 |
-
we widen into the legacy 7-bucket FairFace schema so the
|
| 13 |
-
rest of the app's distribution shape doesn't change.
|
| 14 |
-
|
| 15 |
-
All three are Apache 2.0 and Hugging Face image-classification pipelines.
|
| 16 |
-
|
| 17 |
-
Inputs
|
| 18 |
-
------
|
| 19 |
-
img_rgb : np.ndarray (H, W, 3) uint8
|
| 20 |
-
|
| 21 |
-
Outputs (dict)
|
| 22 |
-
--------------
|
| 23 |
-
age_range, age_estimate (softmax-weighted continuous), age_confidence,
|
| 24 |
-
age_distribution, gender, gender_confidence, ethnicity,
|
| 25 |
-
ethnicity_confidence, ethnicity_distribution.
|
| 26 |
-
|
| 27 |
-
Notes
|
| 28 |
-
-----
|
| 29 |
-
The FairFace age model is a 9-bucket classifier (0-2, 3-9, β¦, 70+),
|
| 30 |
-
which means the argmax bucket midpoint is always one of nine fixed
|
| 31 |
-
numbers (24.5 for 20-29, etc.). To recover a smooth continuous estimate
|
| 32 |
-
we compute the expected value across the full softmax β see
|
| 33 |
-
``_weighted_age_estimate``.
|
| 34 |
-
"""
|
| 35 |
-
|
| 36 |
-
from typing import Any
|
| 37 |
-
|
| 38 |
-
from PIL import Image
|
| 39 |
-
from transformers import pipeline
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
AGE_MODEL_ID = "dima806/fairface_age_image_detection"
|
| 43 |
-
GENDER_MODEL_ID = "dima806/fairface_gender_image_detection"
|
| 44 |
-
RACE_MODEL_ID = "cledoux42/Ethnicity_Test_v003"
|
| 45 |
-
|
| 46 |
-
AGE_LABELS = ["0-2", "3-9", "10-19", "20-29", "30-39", "40-49", "50-59", "60-69", "70+"]
|
| 47 |
-
GENDER_LABELS = ["Male", "Female"]
|
| 48 |
-
# cledoux42 ships 5 classes (african, asian, caucasian, hispanic, indian),
|
| 49 |
-
# but we keep the legacy 7-bucket FairFace label space internally so the
|
| 50 |
-
# downstream distribution dict shape stays stable. Unseen buckets stay 0.
|
| 51 |
-
RACE_LABELS = ["White", "Black", "Latino_Hispanic", "East Asian", "Southeast Asian", "Indian", "Middle Eastern"]
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
class DemographicAnalyzer:
|
| 55 |
-
def __init__(self):
|
| 56 |
-
# Each classifier is a HF image-classification pipeline. They lazy
|
| 57 |
-
# download weights from HF on first instantiation and cache them
|
| 58 |
-
# under /root/.cache/huggingface inside the container.
|
| 59 |
-
self.age_classifier = self._load_classifier(AGE_MODEL_ID)
|
| 60 |
-
self.gender_classifier = self._load_classifier(GENDER_MODEL_ID)
|
| 61 |
-
self.race_classifier = self._load_classifier(RACE_MODEL_ID)
|
| 62 |
-
|
| 63 |
-
@staticmethod
|
| 64 |
-
def _load_classifier(model_id: str):
|
| 65 |
-
"""Build one HF image-classification pipeline, logging on failure.
|
| 66 |
-
|
| 67 |
-
A failed load returns None so the rest of the service continues
|
| 68 |
-
to function and `analyze()` falls back to "unknown" demographics.
|
| 69 |
-
"""
|
| 70 |
-
try:
|
| 71 |
-
return pipeline("image-classification", model=model_id)
|
| 72 |
-
except Exception as exc:
|
| 73 |
-
print(f"[DemographicAnalyzer] Failed to load '{model_id}': {exc}")
|
| 74 |
-
return None
|
| 75 |
-
|
| 76 |
-
def analyze(self, img_rgb) -> dict[str, Any]:
|
| 77 |
-
# Convert the numpy frame to a PIL Image once and reuse it for
|
| 78 |
-
# all three classifier calls.
|
| 79 |
-
pil = Image.fromarray(img_rgb)
|
| 80 |
-
|
| 81 |
-
# top_k=len(labels) so we get the full softmax for each model.
|
| 82 |
-
# We need the full age distribution to compute the weighted
|
| 83 |
-
# expected-value age estimate.
|
| 84 |
-
age_predictions = self._safe_predict(self.age_classifier, pil, top_k=len(AGE_LABELS))
|
| 85 |
-
gender_predictions = self._safe_predict(self.gender_classifier, pil, top_k=2)
|
| 86 |
-
race_predictions = self._safe_predict(self.race_classifier, pil, top_k=7)
|
| 87 |
-
|
| 88 |
-
# If every classifier failed we degrade gracefully with a stub.
|
| 89 |
-
if not age_predictions and not gender_predictions and not race_predictions:
|
| 90 |
-
return {
|
| 91 |
-
"age_range": "unknown",
|
| 92 |
-
"age_estimate": 0.0,
|
| 93 |
-
"age_confidence": 0.0,
|
| 94 |
-
"gender": "unknown",
|
| 95 |
-
"gender_confidence": 0.0,
|
| 96 |
-
"ethnicity": "unknown",
|
| 97 |
-
"ethnicity_confidence": 0.0,
|
| 98 |
-
"age_distribution": {label: 0.0 for label in AGE_LABELS},
|
| 99 |
-
"ethnicity_distribution": {label: 0.0 for label in RACE_LABELS},
|
| 100 |
-
}
|
| 101 |
-
|
| 102 |
-
# HF pipelines return predictions pre-sorted by score descending,
|
| 103 |
-
# so prediction[0] is always the argmax class.
|
| 104 |
-
age_prediction = age_predictions[0] if age_predictions else {"label": "unknown", "score": 0.0}
|
| 105 |
-
gender_prediction = gender_predictions[0] if gender_predictions else {"label": "unknown", "score": 0.0}
|
| 106 |
-
race_prediction = race_predictions[0] if race_predictions else {"label": "unknown", "score": 0.0}
|
| 107 |
-
|
| 108 |
-
# Models occasionally return label aliases ("more than 70" instead
|
| 109 |
-
# of "70+", "African" instead of "Black"). The normalisers map
|
| 110 |
-
# everything back to our canonical schema.
|
| 111 |
-
age_label = self._normalize_age_label(age_prediction["label"])
|
| 112 |
-
gender_label = self._normalize_gender_label(gender_prediction["label"])
|
| 113 |
-
race_label = self._normalize_race_label(race_prediction["label"])
|
| 114 |
-
|
| 115 |
-
return {
|
| 116 |
-
"age_range": age_label,
|
| 117 |
-
"age_estimate": self._weighted_age_estimate(age_predictions),
|
| 118 |
-
"age_confidence": round(float(age_prediction["score"]), 3),
|
| 119 |
-
"gender": gender_label.lower(),
|
| 120 |
-
"gender_confidence": round(float(gender_prediction["score"]), 3),
|
| 121 |
-
"ethnicity": race_label,
|
| 122 |
-
"ethnicity_confidence": round(float(race_prediction["score"]), 3),
|
| 123 |
-
"age_distribution": self._distribution_map(age_predictions, self._normalize_age_label, AGE_LABELS),
|
| 124 |
-
"ethnicity_distribution": self._distribution_map(race_predictions, self._normalize_race_label, RACE_LABELS),
|
| 125 |
-
}
|
| 126 |
-
|
| 127 |
-
@staticmethod
|
| 128 |
-
def _normalize_age_label(label: str) -> str:
|
| 129 |
-
"""Map model output to canonical AGE_LABELS entry."""
|
| 130 |
-
normalized = label.strip().lower()
|
| 131 |
-
if normalized == "more than 70":
|
| 132 |
-
return "70+"
|
| 133 |
-
return AGE_LABELS[AGE_LABELS.index(label)] if label in AGE_LABELS else label
|
| 134 |
-
|
| 135 |
-
@staticmethod
|
| 136 |
-
def _normalize_gender_label(label: str) -> str:
|
| 137 |
-
normalized = label.strip().lower()
|
| 138 |
-
if normalized in {"male", "female"}:
|
| 139 |
-
return normalized.capitalize()
|
| 140 |
-
return label
|
| 141 |
-
|
| 142 |
-
@staticmethod
|
| 143 |
-
def _normalize_race_label(label: str) -> str:
|
| 144 |
-
"""Coalesce cledoux42's 5 classes into our 7-bucket schema."""
|
| 145 |
-
normalized = label.strip().lower().replace("-", "_")
|
| 146 |
-
race_aliases = {
|
| 147 |
-
# Legacy FairFace 7-class labels
|
| 148 |
-
"white": "White",
|
| 149 |
-
"black": "Black",
|
| 150 |
-
"latino_hispanic": "Latino_Hispanic",
|
| 151 |
-
"latino hispanic": "Latino_Hispanic",
|
| 152 |
-
"east asian": "East Asian",
|
| 153 |
-
"southeast asian": "Southeast Asian",
|
| 154 |
-
"indian": "Indian",
|
| 155 |
-
"middle eastern": "Middle Eastern",
|
| 156 |
-
# cledoux42/Ethnicity_Test_v003 5-class labels
|
| 157 |
-
"african": "Black",
|
| 158 |
-
"asian": "East Asian",
|
| 159 |
-
"caucasian": "White",
|
| 160 |
-
"hispanic": "Latino_Hispanic",
|
| 161 |
-
}
|
| 162 |
-
return race_aliases.get(normalized, label)
|
| 163 |
-
|
| 164 |
-
# Midpoint of each FairFace age bucket β used as the per-bucket
|
| 165 |
-
# "value" when we marginalise over the predicted distribution.
|
| 166 |
-
_AGE_MIDPOINTS = {
|
| 167 |
-
"0-2": 1.0,
|
| 168 |
-
"3-9": 6.0,
|
| 169 |
-
"10-19": 14.5,
|
| 170 |
-
"20-29": 24.5,
|
| 171 |
-
"30-39": 34.5,
|
| 172 |
-
"40-49": 44.5,
|
| 173 |
-
"50-59": 54.5,
|
| 174 |
-
"60-69": 64.5,
|
| 175 |
-
"70+": 75.0,
|
| 176 |
-
}
|
| 177 |
-
|
| 178 |
-
@classmethod
|
| 179 |
-
def _weighted_age_estimate(cls, predictions: list[dict]) -> float:
|
| 180 |
-
"""Softmax-weighted expected age across all FairFace buckets.
|
| 181 |
-
|
| 182 |
-
FairFace is a 9-bucket classifier; the argmax always snaps to one
|
| 183 |
-
of nine fixed midpoints (24.5 for 20-29, etc.). Treating its
|
| 184 |
-
softmax as a probability distribution and taking the expected
|
| 185 |
-
value gives a continuous number that moves with confidence
|
| 186 |
-
(23.1 for someone very confidently 20-29, 28.4 if some mass leaks
|
| 187 |
-
into 30-39). Still bounded by bucket midpoints β true per-year
|
| 188 |
-
accuracy would need a regression model.
|
| 189 |
-
"""
|
| 190 |
-
total_weight = 0.0
|
| 191 |
-
weighted_sum = 0.0
|
| 192 |
-
for pred in predictions:
|
| 193 |
-
label = cls._normalize_age_label(pred["label"])
|
| 194 |
-
midpoint = cls._AGE_MIDPOINTS.get(label)
|
| 195 |
-
if midpoint is None:
|
| 196 |
-
continue
|
| 197 |
-
score = float(pred["score"])
|
| 198 |
-
weighted_sum += midpoint * score
|
| 199 |
-
total_weight += score
|
| 200 |
-
if total_weight == 0:
|
| 201 |
-
return 0.0
|
| 202 |
-
return round(weighted_sum / total_weight, 1)
|
| 203 |
-
|
| 204 |
-
@classmethod
|
| 205 |
-
def _distribution_map(cls, predictions, normalizer, all_labels):
|
| 206 |
-
"""Flatten HF predictions into {canonical_label: score} dict.
|
| 207 |
-
|
| 208 |
-
Unseen labels stay at 0.0 so the shape is always all_labels-sized.
|
| 209 |
-
"""
|
| 210 |
-
distribution = {label: 0.0 for label in all_labels}
|
| 211 |
-
for prediction in predictions:
|
| 212 |
-
normalized_label = normalizer(prediction["label"])
|
| 213 |
-
if normalized_label in distribution:
|
| 214 |
-
distribution[normalized_label] = round(float(prediction["score"]), 3)
|
| 215 |
-
return distribution
|
| 216 |
-
|
| 217 |
-
@staticmethod
|
| 218 |
-
def _safe_predict(classifier, image, top_k: int):
|
| 219 |
-
"""Wrap classifier(...) so a single model failure can't bring
|
| 220 |
-
down the whole demographic block."""
|
| 221 |
-
if classifier is None:
|
| 222 |
-
return []
|
| 223 |
-
try:
|
| 224 |
-
return classifier(image, top_k=top_k)
|
| 225 |
-
except Exception as exc:
|
| 226 |
-
print(f"[DemographicAnalyzer] Prediction failed: {exc}")
|
| 227 |
-
return []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
analyzers/ethnicity_analyzer.py
ADDED
|
@@ -0,0 +1,115 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
EthnicityAnalyzer β single-model ethnicity classifier.
|
| 3 |
+
|
| 4 |
+
Model
|
| 5 |
+
-----
|
| 6 |
+
- HF repo : cledoux42/Ethnicity_Test_v003
|
| 7 |
+
- Arch : Vision Transformer
|
| 8 |
+
- Classes : african, asian, caucasian, hispanic, indian (5)
|
| 9 |
+
- Reported : 79.6% accuracy, macro-F1 0.797
|
| 10 |
+
- License : Apache 2.0
|
| 11 |
+
- Source : https://huggingface.co/cledoux42/Ethnicity_Test_v003
|
| 12 |
+
|
| 13 |
+
Inputs
|
| 14 |
+
------
|
| 15 |
+
img_rgb : np.ndarray (H, W, 3) uint8
|
| 16 |
+
|
| 17 |
+
Outputs (dict)
|
| 18 |
+
--------------
|
| 19 |
+
ethnicity : canonical label (legacy 7-bucket schema)
|
| 20 |
+
ethnicity_confidence : argmax softmax score
|
| 21 |
+
ethnicity_distribution : full {label: prob} dict, padded to all 7 buckets
|
| 22 |
+
|
| 23 |
+
Notes
|
| 24 |
+
-----
|
| 25 |
+
Split out from the old DemographicAnalyzer (which also handled age and
|
| 26 |
+
gender via FairFace ViTs). Age and gender now live in
|
| 27 |
+
InsightFaceAnalyzer; this file owns ethnicity exclusively.
|
| 28 |
+
|
| 29 |
+
The model emits 5 classes but we widen to the legacy 7-bucket FairFace
|
| 30 |
+
schema so the rest of the app's distribution shape stays stable.
|
| 31 |
+
Unseen buckets stay at 0.0.
|
| 32 |
+
"""
|
| 33 |
+
|
| 34 |
+
from typing import Any
|
| 35 |
+
|
| 36 |
+
from PIL import Image
|
| 37 |
+
from transformers import pipeline
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
MODEL_ID = "cledoux42/Ethnicity_Test_v003"
|
| 41 |
+
|
| 42 |
+
# Legacy schema preserved from the old DemographicAnalyzer.
|
| 43 |
+
RACE_LABELS = [
|
| 44 |
+
"White", "Black", "Latino_Hispanic", "East Asian",
|
| 45 |
+
"Southeast Asian", "Indian", "Middle Eastern",
|
| 46 |
+
]
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
class EthnicityAnalyzer:
|
| 50 |
+
def __init__(self):
|
| 51 |
+
self.classifier = None
|
| 52 |
+
try:
|
| 53 |
+
self.classifier = pipeline("image-classification", model=MODEL_ID)
|
| 54 |
+
except Exception as exc:
|
| 55 |
+
print(f"[EthnicityAnalyzer] Failed to load {MODEL_ID}: {exc}")
|
| 56 |
+
|
| 57 |
+
def analyze(self, img_rgb) -> dict[str, Any]:
|
| 58 |
+
if self.classifier is None:
|
| 59 |
+
return self._empty_result()
|
| 60 |
+
|
| 61 |
+
try:
|
| 62 |
+
pil = Image.fromarray(img_rgb)
|
| 63 |
+
preds = self.classifier(pil, top_k=7)
|
| 64 |
+
except Exception as exc:
|
| 65 |
+
print(f"[EthnicityAnalyzer] Prediction failed: {exc}")
|
| 66 |
+
return self._empty_result()
|
| 67 |
+
|
| 68 |
+
if not preds:
|
| 69 |
+
return self._empty_result()
|
| 70 |
+
|
| 71 |
+
# Top prediction β canonical label.
|
| 72 |
+
top = preds[0]
|
| 73 |
+
top_label = self._normalize(top["label"])
|
| 74 |
+
|
| 75 |
+
distribution = {label: 0.0 for label in RACE_LABELS}
|
| 76 |
+
for pred in preds:
|
| 77 |
+
label = self._normalize(pred["label"])
|
| 78 |
+
if label in distribution:
|
| 79 |
+
distribution[label] = round(float(pred["score"]), 3)
|
| 80 |
+
|
| 81 |
+
return {
|
| 82 |
+
"ethnicity": top_label,
|
| 83 |
+
"ethnicity_confidence": round(float(top["score"]), 3),
|
| 84 |
+
"ethnicity_distribution": distribution,
|
| 85 |
+
}
|
| 86 |
+
|
| 87 |
+
@staticmethod
|
| 88 |
+
def _normalize(label: str) -> str:
|
| 89 |
+
"""Map model output (5-class) to canonical 7-bucket label."""
|
| 90 |
+
normalized = label.strip().lower().replace("-", "_")
|
| 91 |
+
aliases = {
|
| 92 |
+
# Legacy FairFace 7-class
|
| 93 |
+
"white": "White",
|
| 94 |
+
"black": "Black",
|
| 95 |
+
"latino_hispanic": "Latino_Hispanic",
|
| 96 |
+
"latino hispanic": "Latino_Hispanic",
|
| 97 |
+
"east asian": "East Asian",
|
| 98 |
+
"southeast asian": "Southeast Asian",
|
| 99 |
+
"indian": "Indian",
|
| 100 |
+
"middle eastern": "Middle Eastern",
|
| 101 |
+
# cledoux42 5-class β widen into 7-bucket schema
|
| 102 |
+
"african": "Black",
|
| 103 |
+
"asian": "East Asian",
|
| 104 |
+
"caucasian": "White",
|
| 105 |
+
"hispanic": "Latino_Hispanic",
|
| 106 |
+
}
|
| 107 |
+
return aliases.get(normalized, label)
|
| 108 |
+
|
| 109 |
+
@staticmethod
|
| 110 |
+
def _empty_result() -> dict[str, Any]:
|
| 111 |
+
return {
|
| 112 |
+
"ethnicity": "unknown",
|
| 113 |
+
"ethnicity_confidence": 0.0,
|
| 114 |
+
"ethnicity_distribution": {label: 0.0 for label in RACE_LABELS},
|
| 115 |
+
}
|
analyzers/insightface_analyzer.py
ADDED
|
@@ -0,0 +1,171 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
InsightFaceAnalyzer β detection + age + gender + recognition embedding.
|
| 3 |
+
|
| 4 |
+
Model
|
| 5 |
+
-----
|
| 6 |
+
- Package : `insightface` (https://github.com/deepinsight/insightface)
|
| 7 |
+
- Bundle : buffalo_l (ResNet50@WebFace600K backbone, ONNX)
|
| 8 |
+
- Components : SCRFD-10GF detector, ArcFace 512-d recognition,
|
| 9 |
+
2d106 + 3d68 landmark regressors, age + gender heads
|
| 10 |
+
- Size : ~280 MB (ONNX, mixed FP16/FP32)
|
| 11 |
+
- License : weights research-only; code Apache 2.0
|
| 12 |
+
- Source : https://github.com/deepinsight/insightface/tree/master/python-package
|
| 13 |
+
|
| 14 |
+
Inputs
|
| 15 |
+
------
|
| 16 |
+
img_rgb : np.ndarray (H, W, 3) uint8
|
| 17 |
+
|
| 18 |
+
Outputs (dict)
|
| 19 |
+
--------------
|
| 20 |
+
face_bbox : [x1, y1, x2, y2] in pixel coordinates
|
| 21 |
+
face_confidence : SCRFD detection score
|
| 22 |
+
face_embedding : list[float] of length 512 (ArcFace, L2-normalised)
|
| 23 |
+
age_estimate : float years (regression head, not bucketed)
|
| 24 |
+
age_range : string bucket derived from age_estimate for
|
| 25 |
+
backwards compatibility with the legacy UI
|
| 26 |
+
gender : "male" | "female"
|
| 27 |
+
gender_confidence : 1.0 by default (InsightFace doesn't expose a
|
| 28 |
+
gender softmax score; the head is argmax-only)
|
| 29 |
+
_insight_landmarks_2d : list of (x, y) tuples β 106 points (internal)
|
| 30 |
+
|
| 31 |
+
Accuracy
|
| 32 |
+
--------
|
| 33 |
+
- Recognition (ArcFace via buffalo_l): 99.83% LFW, 96.21% IJB-B FAR=1e-4.
|
| 34 |
+
- Age / gender heads are widely used but lack a clean published metric.
|
| 35 |
+
In practice age MAE is ~5 years and gender ~94-96%.
|
| 36 |
+
|
| 37 |
+
Notes
|
| 38 |
+
-----
|
| 39 |
+
We run the bundle once per image and expose only the highest-confidence
|
| 40 |
+
face when multiple are detected β the rest of the pipeline assumes a
|
| 41 |
+
single subject.
|
| 42 |
+
"""
|
| 43 |
+
|
| 44 |
+
import os
|
| 45 |
+
from typing import Any
|
| 46 |
+
|
| 47 |
+
import numpy as np
|
| 48 |
+
|
| 49 |
+
# insightface is a relatively heavy import; deferred so the module can
|
| 50 |
+
# at least load when the package isn't installed.
|
| 51 |
+
try:
|
| 52 |
+
from insightface.app import FaceAnalysis
|
| 53 |
+
HAS_INSIGHTFACE = True
|
| 54 |
+
except ImportError:
|
| 55 |
+
HAS_INSIGHTFACE = False
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
MODEL_NAME = "buffalo_l"
|
| 59 |
+
|
| 60 |
+
# Age buckets used by the legacy UI. We derive these from the regression
|
| 61 |
+
# output so existing screens keep working.
|
| 62 |
+
AGE_BUCKETS = [
|
| 63 |
+
(0, 3, "0-2"), (3, 10, "3-9"), (10, 20, "10-19"),
|
| 64 |
+
(20, 30, "20-29"), (30, 40, "30-39"), (40, 50, "40-49"),
|
| 65 |
+
(50, 60, "50-59"), (60, 70, "60-69"), (70, 200, "70+"),
|
| 66 |
+
]
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
class InsightFaceAnalyzer:
|
| 70 |
+
def __init__(self):
|
| 71 |
+
self.app = None
|
| 72 |
+
if not HAS_INSIGHTFACE:
|
| 73 |
+
print(
|
| 74 |
+
"[InsightFaceAnalyzer] insightface package not installed; "
|
| 75 |
+
"detection, age, gender, and recognition will degrade to 'unknown'."
|
| 76 |
+
)
|
| 77 |
+
return
|
| 78 |
+
|
| 79 |
+
try:
|
| 80 |
+
# Buffalo_L bundle auto-resolves under ~/.insightface/models/.
|
| 81 |
+
# CPUExecutionProvider is the right default for HF Spaces;
|
| 82 |
+
# ctx_id=0 + 'CUDAExecutionProvider' would be the GPU path.
|
| 83 |
+
self.app = FaceAnalysis(
|
| 84 |
+
name=MODEL_NAME,
|
| 85 |
+
providers=["CPUExecutionProvider"],
|
| 86 |
+
)
|
| 87 |
+
# det_size=(640, 640) is the canonical SCRFD input. Smaller
|
| 88 |
+
# speeds inference but loses small faces.
|
| 89 |
+
self.app.prepare(ctx_id=-1, det_size=(640, 640))
|
| 90 |
+
except Exception as exc:
|
| 91 |
+
print(f"[InsightFaceAnalyzer] Failed to load {MODEL_NAME}: {exc}")
|
| 92 |
+
self.app = None
|
| 93 |
+
|
| 94 |
+
def analyze(self, img_rgb: np.ndarray) -> dict[str, Any]:
|
| 95 |
+
if self.app is None:
|
| 96 |
+
return self._empty_result()
|
| 97 |
+
|
| 98 |
+
try:
|
| 99 |
+
# InsightFace expects BGR, OpenCV-native order.
|
| 100 |
+
img_bgr = img_rgb[..., ::-1]
|
| 101 |
+
faces = self.app.get(img_bgr)
|
| 102 |
+
except Exception as exc:
|
| 103 |
+
print(f"[InsightFaceAnalyzer] Inference failed: {exc}")
|
| 104 |
+
return self._empty_result()
|
| 105 |
+
|
| 106 |
+
if not faces:
|
| 107 |
+
return self._empty_result()
|
| 108 |
+
|
| 109 |
+
# If multiple faces, take the highest-confidence one. The rest of
|
| 110 |
+
# the pipeline assumes a single subject.
|
| 111 |
+
face = max(faces, key=lambda f: float(f.det_score))
|
| 112 |
+
|
| 113 |
+
# Bounding box is float32 in [x1, y1, x2, y2] image-pixel space.
|
| 114 |
+
bbox = [float(v) for v in face.bbox.tolist()]
|
| 115 |
+
|
| 116 |
+
# Recognition embedding: 512-d, L2-normalised by InsightFace.
|
| 117 |
+
# Cast to plain list[float] for clean JSON.
|
| 118 |
+
embedding = (
|
| 119 |
+
[float(v) for v in face.normed_embedding.tolist()]
|
| 120 |
+
if getattr(face, "normed_embedding", None) is not None
|
| 121 |
+
else None
|
| 122 |
+
)
|
| 123 |
+
|
| 124 |
+
# Age head is a single float (years). Round for display.
|
| 125 |
+
age = float(getattr(face, "age", 0.0))
|
| 126 |
+
|
| 127 |
+
# Gender is exposed as 0 (female) / 1 (male) on Face objects.
|
| 128 |
+
# InsightFace doesn't surface a softmax probability β we report
|
| 129 |
+
# confidence 1.0 to indicate "argmax, no soft signal".
|
| 130 |
+
gender_idx = int(getattr(face, "gender", -1))
|
| 131 |
+
gender = "male" if gender_idx == 1 else "female" if gender_idx == 0 else "unknown"
|
| 132 |
+
|
| 133 |
+
return {
|
| 134 |
+
"face_bbox": bbox,
|
| 135 |
+
"face_confidence": round(float(face.det_score), 3),
|
| 136 |
+
"face_embedding": embedding,
|
| 137 |
+
"age_estimate": round(age, 1),
|
| 138 |
+
"age_range": self._bucket_age(age),
|
| 139 |
+
"age_confidence": 1.0,
|
| 140 |
+
"gender": gender,
|
| 141 |
+
"gender_confidence": 1.0,
|
| 142 |
+
# 106 2D landmarks (forehead, jaw, brows, eyes, nose, lips).
|
| 143 |
+
# Underscore-prefixed β stripped from JSON, available to
|
| 144 |
+
# downstream analyzers that want a tighter face crop.
|
| 145 |
+
"_insight_landmarks_2d": (
|
| 146 |
+
[(float(p[0]), float(p[1])) for p in face.landmark_2d_106.tolist()]
|
| 147 |
+
if getattr(face, "landmark_2d_106", None) is not None
|
| 148 |
+
else None
|
| 149 |
+
),
|
| 150 |
+
}
|
| 151 |
+
|
| 152 |
+
@staticmethod
|
| 153 |
+
def _bucket_age(age: float) -> str:
|
| 154 |
+
for lo, hi, label in AGE_BUCKETS:
|
| 155 |
+
if lo <= age < hi:
|
| 156 |
+
return label
|
| 157 |
+
return "unknown"
|
| 158 |
+
|
| 159 |
+
@staticmethod
|
| 160 |
+
def _empty_result() -> dict[str, Any]:
|
| 161 |
+
return {
|
| 162 |
+
"face_bbox": None,
|
| 163 |
+
"face_confidence": 0.0,
|
| 164 |
+
"face_embedding": None,
|
| 165 |
+
"age_estimate": 0.0,
|
| 166 |
+
"age_range": "unknown",
|
| 167 |
+
"age_confidence": 0.0,
|
| 168 |
+
"gender": "unknown",
|
| 169 |
+
"gender_confidence": 0.0,
|
| 170 |
+
"_insight_landmarks_2d": None,
|
| 171 |
+
}
|
analyzers/landmark_analyzer.py
CHANGED
|
@@ -11,6 +11,9 @@ Model
|
|
| 11 |
Inputs
|
| 12 |
------
|
| 13 |
img_rgb : np.ndarray (H, W, 3) uint8, RGB order.
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
Outputs (dict)
|
| 16 |
--------------
|
|
|
|
| 11 |
Inputs
|
| 12 |
------
|
| 13 |
img_rgb : np.ndarray (H, W, 3) uint8, RGB order.
|
| 14 |
+
Always receives the full photo (not a face crop) β MediaPipe
|
| 15 |
+
has its own face detector and works best with the full
|
| 16 |
+
field of view to maintain z-coordinate consistency.
|
| 17 |
|
| 18 |
Outputs (dict)
|
| 19 |
--------------
|
analyzers/parsing_analyzer.py
CHANGED
|
@@ -15,7 +15,12 @@ Model
|
|
| 15 |
|
| 16 |
Inputs
|
| 17 |
------
|
| 18 |
-
img_rgb : np.ndarray (H, W, 3) uint8
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
Outputs (dict)
|
| 21 |
--------------
|
|
|
|
| 15 |
|
| 16 |
Inputs
|
| 17 |
------
|
| 18 |
+
img_rgb : np.ndarray (H, W, 3) uint8.
|
| 19 |
+
Since `app.py` v3.0, this is typically a face-cropped image
|
| 20 |
+
(produced by `_crop_to_face` using the InsightFace bbox)
|
| 21 |
+
rather than the full photo. SegFormer behaves much more
|
| 22 |
+
consistently when the face fills most of the input β the
|
| 23 |
+
masks become tighter and skin-stat estimates less noisy.
|
| 24 |
|
| 25 |
Outputs (dict)
|
| 26 |
--------------
|
app.py
CHANGED
|
@@ -2,38 +2,57 @@
|
|
| 2 |
HCP Face Analysis Microservice
|
| 3 |
==============================
|
| 4 |
|
| 5 |
-
FastAPI service that runs
|
| 6 |
-
and merges their outputs into one
|
|
|
|
|
|
|
| 7 |
|
| 8 |
Pipeline (in execution order)
|
| 9 |
-----------------------------
|
| 10 |
-
1.
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
6.
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
Endpoints
|
| 39 |
---------
|
|
@@ -42,15 +61,14 @@ GET /health liveness check
|
|
| 42 |
POST /analyze multipart file upload
|
| 43 |
POST /analyze-base64 JSON {"image": "<base64>"}
|
| 44 |
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
on the Hugging Face Spaces free tier.
|
| 48 |
"""
|
| 49 |
|
| 50 |
import os
|
| 51 |
-
# hf_transfer
|
| 52 |
-
#
|
| 53 |
-
#
|
| 54 |
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
|
| 55 |
os.environ["HF_HUB_DOWNLOAD_TIMEOUT"] = "60"
|
| 56 |
|
|
@@ -64,34 +82,41 @@ from fastapi.middleware.cors import CORSMiddleware
|
|
| 64 |
from PIL import Image
|
| 65 |
|
| 66 |
from analyzers.landmark_analyzer import LandmarkAnalyzer
|
| 67 |
-
from analyzers.
|
| 68 |
from analyzers.parsing_analyzer import ParsingAnalyzer
|
| 69 |
from analyzers.emotion_analyzer import EmotionAnalyzer
|
| 70 |
from analyzers.color_analyzer import ColorAnalyzer
|
| 71 |
from analyzers.obstruction_analyzer import ObstructionAnalyzer
|
| 72 |
from analyzers.hair_type_analyzer import HairTypeAnalyzer
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
logging.basicConfig(level=logging.INFO)
|
| 75 |
logger = logging.getLogger(__name__)
|
| 76 |
|
| 77 |
-
app = FastAPI(title="HCP Face Analysis Service", version="
|
| 78 |
|
| 79 |
app.add_middleware(
|
| 80 |
CORSMiddleware,
|
| 81 |
-
allow_origins=["*"], # Restrict to your domain in production
|
| 82 |
allow_credentials=True,
|
| 83 |
allow_methods=["*"],
|
| 84 |
allow_headers=["*"],
|
| 85 |
)
|
| 86 |
|
| 87 |
-
#
|
|
|
|
|
|
|
| 88 |
landmark_analyzer: Optional[LandmarkAnalyzer] = None
|
| 89 |
-
|
| 90 |
parsing_analyzer: Optional[ParsingAnalyzer] = None
|
| 91 |
emotion_analyzer: Optional[EmotionAnalyzer] = None
|
| 92 |
color_analyzer: Optional[ColorAnalyzer] = None
|
| 93 |
obstruction_analyzer: Optional[ObstructionAnalyzer] = None
|
| 94 |
hair_type_analyzer: Optional[HairTypeAnalyzer] = None
|
|
|
|
|
|
|
| 95 |
|
| 96 |
|
| 97 |
def _to_json_safe(value):
|
|
@@ -101,8 +126,6 @@ def _to_json_safe(value):
|
|
| 101 |
or boolean mask logic). FastAPI's default JSON encoder doesn't
|
| 102 |
handle those, so we normalise everything here before returning.
|
| 103 |
"""
|
| 104 |
-
# Numpy first β these checks would otherwise be caught by isinstance
|
| 105 |
-
# for dict/list because numpy.generic types are duck-typed.
|
| 106 |
if isinstance(value, (np.ndarray,)):
|
| 107 |
return value.tolist()
|
| 108 |
if isinstance(value, (np.integer, np.floating)):
|
|
@@ -111,7 +134,6 @@ def _to_json_safe(value):
|
|
| 111 |
return bool(value)
|
| 112 |
if isinstance(value, np.generic):
|
| 113 |
return value.item()
|
| 114 |
-
# Recurse into nested containers.
|
| 115 |
if isinstance(value, dict):
|
| 116 |
return {str(k): _to_json_safe(v) for k, v in value.items()}
|
| 117 |
if isinstance(value, (list, tuple, set)):
|
|
@@ -126,17 +148,22 @@ def get_analyzers():
|
|
| 126 |
requests. First request pays the full model-load cost; subsequent
|
| 127 |
requests are warm.
|
| 128 |
"""
|
| 129 |
-
global landmark_analyzer,
|
| 130 |
global parsing_analyzer, emotion_analyzer, color_analyzer
|
| 131 |
global obstruction_analyzer, hair_type_analyzer
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 132 |
|
| 133 |
if landmark_analyzer is None:
|
| 134 |
logger.info("Loading MediaPipe Face Landmarker...")
|
| 135 |
landmark_analyzer = LandmarkAnalyzer()
|
| 136 |
|
| 137 |
-
if
|
| 138 |
-
logger.info("Loading
|
| 139 |
-
|
| 140 |
|
| 141 |
if parsing_analyzer is None:
|
| 142 |
logger.info("Loading SegFormer face parser...")
|
|
@@ -157,23 +184,156 @@ def get_analyzers():
|
|
| 157 |
logger.info("Loading hair type classifier...")
|
| 158 |
hair_type_analyzer = HairTypeAnalyzer()
|
| 159 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 160 |
return (
|
|
|
|
| 161 |
landmark_analyzer,
|
| 162 |
-
|
| 163 |
parsing_analyzer,
|
| 164 |
emotion_analyzer,
|
| 165 |
color_analyzer,
|
| 166 |
obstruction_analyzer,
|
| 167 |
hair_type_analyzer,
|
|
|
|
|
|
|
| 168 |
)
|
| 169 |
|
| 170 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 171 |
@app.get("/")
|
| 172 |
async def root():
|
| 173 |
"""Service banner β confirms the server is reachable and which version."""
|
| 174 |
return {
|
| 175 |
"name": "HCP Face Analysis Service",
|
| 176 |
-
"version": "
|
| 177 |
"status": "running",
|
| 178 |
"endpoints": {
|
| 179 |
"health": "/health",
|
|
@@ -193,72 +353,15 @@ async def health():
|
|
| 193 |
async def analyze_face(file: UploadFile = File(...)):
|
| 194 |
"""Multipart endpoint for direct uploads.
|
| 195 |
|
| 196 |
-
Runs
|
| 197 |
-
See `analyze_face_base64` for the JSON-body variant the
|
| 198 |
-
server calls.
|
| 199 |
"""
|
| 200 |
try:
|
| 201 |
-
# Decode the upload into an RGB numpy array. All analyzers
|
| 202 |
-
# work in RGB; we don't actually need BGR but keeping it as a
|
| 203 |
-
# local in case a future analyzer wants the OpenCV-native order.
|
| 204 |
contents = await file.read()
|
| 205 |
image = Image.open(io.BytesIO(contents)).convert("RGB")
|
| 206 |
img_array = np.array(image)
|
| 207 |
-
|
| 208 |
-
(
|
| 209 |
-
landmarks,
|
| 210 |
-
demographics,
|
| 211 |
-
parsing,
|
| 212 |
-
emotions,
|
| 213 |
-
colors,
|
| 214 |
-
obstructions,
|
| 215 |
-
hair_types,
|
| 216 |
-
) = get_analyzers()
|
| 217 |
-
|
| 218 |
-
results = {}
|
| 219 |
-
|
| 220 |
-
# Step 1: MediaPipe Landmarks β all geometric features + blendshapes.
|
| 221 |
-
logger.info("Running landmark analysis...")
|
| 222 |
-
landmark_results = landmarks.analyze(img_array)
|
| 223 |
-
results.update(landmark_results)
|
| 224 |
-
|
| 225 |
-
# Step 2: FairFace + Ethnicity ViT β demographics.
|
| 226 |
-
logger.info("Running demographic analysis...")
|
| 227 |
-
demo_results = demographics.analyze(img_array)
|
| 228 |
-
results.update(demo_results)
|
| 229 |
-
|
| 230 |
-
# Step 3: SegFormer-B5 human parsing β masks + hair length + skin stats.
|
| 231 |
-
logger.info("Running face parsing...")
|
| 232 |
-
parse_results = parsing.analyze(img_array)
|
| 233 |
-
results.update(parse_results)
|
| 234 |
-
|
| 235 |
-
# Step 4: HSEmotion β 8-class emotion + valence/arousal/mood.
|
| 236 |
-
logger.info("Running emotion analysis...")
|
| 237 |
-
emo_results = emotions.analyze(img_array)
|
| 238 |
-
results.update(emo_results)
|
| 239 |
-
|
| 240 |
-
# Step 5: Pixel color analysis. Uses the face/hair masks from step 3
|
| 241 |
-
# and MediaPipe lip/iris landmarks from step 1.
|
| 242 |
-
logger.info("Running color analysis...")
|
| 243 |
-
color_results = colors.analyze(
|
| 244 |
-
img_array,
|
| 245 |
-
skin_mask=parse_results.get("_skin_mask"),
|
| 246 |
-
hair_mask=parse_results.get("_hair_mask"),
|
| 247 |
-
landmarks=landmark_results.get("_raw_landmarks"),
|
| 248 |
-
)
|
| 249 |
-
results.update(color_results)
|
| 250 |
-
|
| 251 |
-
# Step 6: ObstructionViT β glasses / sunglasses / mask flags.
|
| 252 |
-
logger.info("Running obstruction analysis...")
|
| 253 |
-
results.update(obstructions.analyze(img_array))
|
| 254 |
-
|
| 255 |
-
# Step 7: HairTypeViT β curly/dreadlocks/kinky/straight/wavy.
|
| 256 |
-
logger.info("Running hair-type analysis...")
|
| 257 |
-
results.update(hair_types.analyze(img_array))
|
| 258 |
-
|
| 259 |
-
# Remove internal fields (prefixed with underscore)
|
| 260 |
-
results = {k: v for k, v in results.items() if not k.startswith("_")}
|
| 261 |
-
|
| 262 |
return {"success": True, "data": _to_json_safe(results)}
|
| 263 |
|
| 264 |
except Exception as e:
|
|
@@ -281,56 +384,14 @@ async def analyze_face_base64(body: dict):
|
|
| 281 |
if not image_b64:
|
| 282 |
raise HTTPException(status_code=400, detail="No image data provided")
|
| 283 |
|
| 284 |
-
# Strip
|
| 285 |
if "," in image_b64:
|
| 286 |
image_b64 = image_b64.split(",", 1)[1]
|
| 287 |
|
| 288 |
image_bytes = base64.b64decode(image_b64)
|
| 289 |
image = Image.open(io.BytesIO(image_bytes)).convert("RGB")
|
| 290 |
img_array = np.array(image)
|
| 291 |
-
|
| 292 |
-
(
|
| 293 |
-
landmarks,
|
| 294 |
-
demographics,
|
| 295 |
-
parsing,
|
| 296 |
-
emotions,
|
| 297 |
-
colors,
|
| 298 |
-
obstructions,
|
| 299 |
-
hair_types,
|
| 300 |
-
) = get_analyzers()
|
| 301 |
-
|
| 302 |
-
results = {}
|
| 303 |
-
|
| 304 |
-
# Same seven-step pipeline as /analyze. Kept inline (rather
|
| 305 |
-
# than factored out) so the per-step `logger.info` cadence and
|
| 306 |
-
# ordering stay obvious when reading either endpoint top-down.
|
| 307 |
-
landmark_results = landmarks.analyze(img_array)
|
| 308 |
-
results.update(landmark_results)
|
| 309 |
-
|
| 310 |
-
demo_results = demographics.analyze(img_array)
|
| 311 |
-
results.update(demo_results)
|
| 312 |
-
|
| 313 |
-
parse_results = parsing.analyze(img_array)
|
| 314 |
-
results.update(parse_results)
|
| 315 |
-
|
| 316 |
-
emo_results = emotions.analyze(img_array)
|
| 317 |
-
results.update(emo_results)
|
| 318 |
-
|
| 319 |
-
color_results = colors.analyze(
|
| 320 |
-
img_array,
|
| 321 |
-
skin_mask=parse_results.get("_skin_mask"),
|
| 322 |
-
hair_mask=parse_results.get("_hair_mask"),
|
| 323 |
-
landmarks=landmark_results.get("_raw_landmarks"),
|
| 324 |
-
)
|
| 325 |
-
results.update(color_results)
|
| 326 |
-
|
| 327 |
-
results.update(obstructions.analyze(img_array))
|
| 328 |
-
results.update(hair_types.analyze(img_array))
|
| 329 |
-
|
| 330 |
-
# Drop internal/scratch fields (leading underscore) before
|
| 331 |
-
# returning. Keeps masks and raw landmark lists out of the JSON.
|
| 332 |
-
results = {k: v for k, v in results.items() if not k.startswith("_")}
|
| 333 |
-
|
| 334 |
return {"success": True, "data": _to_json_safe(results)}
|
| 335 |
|
| 336 |
except HTTPException:
|
|
|
|
| 2 |
HCP Face Analysis Microservice
|
| 3 |
==============================
|
| 4 |
|
| 5 |
+
FastAPI service that runs nine specialized analyzers over a single photo
|
| 6 |
+
and merges their outputs into one facial-attribute dictionary, including
|
| 7 |
+
a face-recognition embedding for cross-photo grouping and a numeric
|
| 8 |
+
"chopped score" aesthetic rating.
|
| 9 |
|
| 10 |
Pipeline (in execution order)
|
| 11 |
-----------------------------
|
| 12 |
+
1. InsightFaceAnalyzer InsightFace buffalo_l (ONNX). SCRFD
|
| 13 |
+
detection + ArcFace 512-d embedding +
|
| 14 |
+
age regression + gender + 106 landmarks.
|
| 15 |
+
Replaces the previous three FairFace ViTs
|
| 16 |
+
and adds face matching as a new capability.
|
| 17 |
+
|
| 18 |
+
2. LandmarkAnalyzer MediaPipe Face Landmarker. 478 3D
|
| 19 |
+
landmarks + 52 ARKit blendshapes β
|
| 20 |
+
geometric features, smiling, mouth_open.
|
| 21 |
+
|
| 22 |
+
3. EthnicityAnalyzer cledoux42/Ethnicity_Test_v003 ViT.
|
| 23 |
+
5-class ethnicity widened to a 7-bucket
|
| 24 |
+
schema for legacy compatibility.
|
| 25 |
+
|
| 26 |
+
4. ParsingAnalyzer SegFormer-B5 human parsing. Now receives
|
| 27 |
+
a face-cropped image (smaller, cleaner).
|
| 28 |
+
Emits face/hair masks + hair length +
|
| 29 |
+
hat detection + OpenCV-derived skin stats.
|
| 30 |
+
|
| 31 |
+
5. EmotionAnalyzer HSEmotion EfficientNet-B0. 8-class
|
| 32 |
+
emotion + valence/arousal/mood.
|
| 33 |
+
|
| 34 |
+
6. ColorAnalyzer Pure OpenCV LAB/HSV statistics. Uses
|
| 35 |
+
SegFormer masks + MediaPipe lip/iris
|
| 36 |
+
landmarks. No ML model.
|
| 37 |
+
|
| 38 |
+
7. ObstructionAnalyzer dima806 ViT-B/16. Glasses, sunglasses,
|
| 39 |
+
mask. ~99% precision on each.
|
| 40 |
+
|
| 41 |
+
8. HairTypeAnalyzer dima806 ViT-B/16. Curly/dreadlocks/kinky/
|
| 42 |
+
straight/wavy. ~93% accuracy.
|
| 43 |
+
|
| 44 |
+
9. BeautyAnalyzer Optional. ResNet-50 trained on
|
| 45 |
+
SCUT-FBP5500 (see training/beauty/).
|
| 46 |
+
Outputs a 1.0β5.0 beauty score plus a
|
| 47 |
+
0β100 normalised version. Falls back to
|
| 48 |
+
None when no weights are loaded β the
|
| 49 |
+
AestheticAnalyzer then uses rule-based
|
| 50 |
+
scoring only.
|
| 51 |
+
|
| 52 |
+
10. AestheticAnalyzer Pure-Python aggregator. Reads the merged
|
| 53 |
+
dict from analyzers 1β9 and produces the
|
| 54 |
+
final `chopped_score` (0β100, higher =
|
| 55 |
+
more chopped) and a per-factor breakdown.
|
| 56 |
|
| 57 |
Endpoints
|
| 58 |
---------
|
|
|
|
| 61 |
POST /analyze multipart file upload
|
| 62 |
POST /analyze-base64 JSON {"image": "<base64>"}
|
| 63 |
|
| 64 |
+
All analyzers are lazily instantiated on first request to keep
|
| 65 |
+
cold-start latency manageable on the Hugging Face Spaces free tier.
|
|
|
|
| 66 |
"""
|
| 67 |
|
| 68 |
import os
|
| 69 |
+
# hf_transfer makes initial model downloads from the HF Hub much faster.
|
| 70 |
+
# The default HF_HUB_DOWNLOAD_TIMEOUT (10 s) is too short for the larger
|
| 71 |
+
# ViT checkpoints on a cold start.
|
| 72 |
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
|
| 73 |
os.environ["HF_HUB_DOWNLOAD_TIMEOUT"] = "60"
|
| 74 |
|
|
|
|
| 82 |
from PIL import Image
|
| 83 |
|
| 84 |
from analyzers.landmark_analyzer import LandmarkAnalyzer
|
| 85 |
+
from analyzers.ethnicity_analyzer import EthnicityAnalyzer
|
| 86 |
from analyzers.parsing_analyzer import ParsingAnalyzer
|
| 87 |
from analyzers.emotion_analyzer import EmotionAnalyzer
|
| 88 |
from analyzers.color_analyzer import ColorAnalyzer
|
| 89 |
from analyzers.obstruction_analyzer import ObstructionAnalyzer
|
| 90 |
from analyzers.hair_type_analyzer import HairTypeAnalyzer
|
| 91 |
+
from analyzers.insightface_analyzer import InsightFaceAnalyzer
|
| 92 |
+
from analyzers.beauty_analyzer import BeautyAnalyzer
|
| 93 |
+
from analyzers.aesthetic_analyzer import AestheticAnalyzer
|
| 94 |
|
| 95 |
logging.basicConfig(level=logging.INFO)
|
| 96 |
logger = logging.getLogger(__name__)
|
| 97 |
|
| 98 |
+
app = FastAPI(title="HCP Face Analysis Service", version="3.0.0")
|
| 99 |
|
| 100 |
app.add_middleware(
|
| 101 |
CORSMiddleware,
|
| 102 |
+
allow_origins=["*"], # Restrict to your domain in production.
|
| 103 |
allow_credentials=True,
|
| 104 |
allow_methods=["*"],
|
| 105 |
allow_headers=["*"],
|
| 106 |
)
|
| 107 |
|
| 108 |
+
# Lazy slots, one per analyzer. The first request pays the full
|
| 109 |
+
# model-load cost; subsequent requests are warm.
|
| 110 |
+
insightface_analyzer: Optional[InsightFaceAnalyzer] = None
|
| 111 |
landmark_analyzer: Optional[LandmarkAnalyzer] = None
|
| 112 |
+
ethnicity_analyzer: Optional[EthnicityAnalyzer] = None
|
| 113 |
parsing_analyzer: Optional[ParsingAnalyzer] = None
|
| 114 |
emotion_analyzer: Optional[EmotionAnalyzer] = None
|
| 115 |
color_analyzer: Optional[ColorAnalyzer] = None
|
| 116 |
obstruction_analyzer: Optional[ObstructionAnalyzer] = None
|
| 117 |
hair_type_analyzer: Optional[HairTypeAnalyzer] = None
|
| 118 |
+
beauty_analyzer: Optional[BeautyAnalyzer] = None
|
| 119 |
+
aesthetic_analyzer: Optional[AestheticAnalyzer] = None
|
| 120 |
|
| 121 |
|
| 122 |
def _to_json_safe(value):
|
|
|
|
| 126 |
or boolean mask logic). FastAPI's default JSON encoder doesn't
|
| 127 |
handle those, so we normalise everything here before returning.
|
| 128 |
"""
|
|
|
|
|
|
|
| 129 |
if isinstance(value, (np.ndarray,)):
|
| 130 |
return value.tolist()
|
| 131 |
if isinstance(value, (np.integer, np.floating)):
|
|
|
|
| 134 |
return bool(value)
|
| 135 |
if isinstance(value, np.generic):
|
| 136 |
return value.item()
|
|
|
|
| 137 |
if isinstance(value, dict):
|
| 138 |
return {str(k): _to_json_safe(v) for k, v in value.items()}
|
| 139 |
if isinstance(value, (list, tuple, set)):
|
|
|
|
| 148 |
requests. First request pays the full model-load cost; subsequent
|
| 149 |
requests are warm.
|
| 150 |
"""
|
| 151 |
+
global insightface_analyzer, landmark_analyzer, ethnicity_analyzer
|
| 152 |
global parsing_analyzer, emotion_analyzer, color_analyzer
|
| 153 |
global obstruction_analyzer, hair_type_analyzer
|
| 154 |
+
global beauty_analyzer, aesthetic_analyzer
|
| 155 |
+
|
| 156 |
+
if insightface_analyzer is None:
|
| 157 |
+
logger.info("Loading InsightFace buffalo_l bundle...")
|
| 158 |
+
insightface_analyzer = InsightFaceAnalyzer()
|
| 159 |
|
| 160 |
if landmark_analyzer is None:
|
| 161 |
logger.info("Loading MediaPipe Face Landmarker...")
|
| 162 |
landmark_analyzer = LandmarkAnalyzer()
|
| 163 |
|
| 164 |
+
if ethnicity_analyzer is None:
|
| 165 |
+
logger.info("Loading Ethnicity classifier...")
|
| 166 |
+
ethnicity_analyzer = EthnicityAnalyzer()
|
| 167 |
|
| 168 |
if parsing_analyzer is None:
|
| 169 |
logger.info("Loading SegFormer face parser...")
|
|
|
|
| 184 |
logger.info("Loading hair type classifier...")
|
| 185 |
hair_type_analyzer = HairTypeAnalyzer()
|
| 186 |
|
| 187 |
+
if beauty_analyzer is None:
|
| 188 |
+
logger.info("Loading beauty regressor (or no-op if untrained)...")
|
| 189 |
+
beauty_analyzer = BeautyAnalyzer()
|
| 190 |
+
|
| 191 |
+
if aesthetic_analyzer is None:
|
| 192 |
+
aesthetic_analyzer = AestheticAnalyzer()
|
| 193 |
+
|
| 194 |
return (
|
| 195 |
+
insightface_analyzer,
|
| 196 |
landmark_analyzer,
|
| 197 |
+
ethnicity_analyzer,
|
| 198 |
parsing_analyzer,
|
| 199 |
emotion_analyzer,
|
| 200 |
color_analyzer,
|
| 201 |
obstruction_analyzer,
|
| 202 |
hair_type_analyzer,
|
| 203 |
+
beauty_analyzer,
|
| 204 |
+
aesthetic_analyzer,
|
| 205 |
)
|
| 206 |
|
| 207 |
|
| 208 |
+
def _crop_to_face(img_rgb: np.ndarray, bbox, padding: float = 0.4) -> np.ndarray:
|
| 209 |
+
"""Crop the image to a face-centred rectangle with extra context.
|
| 210 |
+
|
| 211 |
+
SegFormer and the ViT classifiers tend to do better with the face
|
| 212 |
+
occupying a large fraction of the input. We pad the InsightFace
|
| 213 |
+
bbox by `padding` (fraction of bbox size) so context like ears,
|
| 214 |
+
hair, and the top of the shoulders is preserved.
|
| 215 |
+
|
| 216 |
+
Returns the full image unchanged if bbox is None, malformed, or
|
| 217 |
+
the resulting crop would be degenerate.
|
| 218 |
+
"""
|
| 219 |
+
if bbox is None or len(bbox) != 4:
|
| 220 |
+
return img_rgb
|
| 221 |
+
h, w = img_rgb.shape[:2]
|
| 222 |
+
try:
|
| 223 |
+
x1, y1, x2, y2 = bbox
|
| 224 |
+
bw = max(1.0, x2 - x1)
|
| 225 |
+
bh = max(1.0, y2 - y1)
|
| 226 |
+
pad_x = bw * padding
|
| 227 |
+
pad_y = bh * padding
|
| 228 |
+
cx1 = max(0, int(x1 - pad_x))
|
| 229 |
+
cy1 = max(0, int(y1 - pad_y))
|
| 230 |
+
cx2 = min(w, int(x2 + pad_x))
|
| 231 |
+
cy2 = min(h, int(y2 + pad_y))
|
| 232 |
+
if cx2 - cx1 < 32 or cy2 - cy1 < 32:
|
| 233 |
+
return img_rgb
|
| 234 |
+
return img_rgb[cy1:cy2, cx1:cx2]
|
| 235 |
+
except Exception:
|
| 236 |
+
return img_rgb
|
| 237 |
+
|
| 238 |
+
|
| 239 |
+
def _run_pipeline(img_array: np.ndarray) -> dict:
|
| 240 |
+
"""Run all ten analyzers against `img_array` and return the merged dict.
|
| 241 |
+
|
| 242 |
+
Shared by /analyze and /analyze-base64. Kept as a function rather
|
| 243 |
+
than inlined twice so the per-step ordering is the single source
|
| 244 |
+
of truth.
|
| 245 |
+
"""
|
| 246 |
+
(
|
| 247 |
+
insight,
|
| 248 |
+
landmarks,
|
| 249 |
+
ethnicities,
|
| 250 |
+
parsing,
|
| 251 |
+
emotions,
|
| 252 |
+
colors,
|
| 253 |
+
obstructions,
|
| 254 |
+
hair_types,
|
| 255 |
+
beauty,
|
| 256 |
+
aesthetics,
|
| 257 |
+
) = get_analyzers()
|
| 258 |
+
|
| 259 |
+
results: dict = {}
|
| 260 |
+
|
| 261 |
+
# Step 1: InsightFace detection + age + gender + recognition embedding.
|
| 262 |
+
logger.info("Running InsightFace analysis...")
|
| 263 |
+
insight_results = insight.analyze(img_array)
|
| 264 |
+
results.update(insight_results)
|
| 265 |
+
|
| 266 |
+
# Compute a face crop once and pass it to every downstream analyzer
|
| 267 |
+
# that benefits from it (parsing, ethnicity, obstruction, hair type,
|
| 268 |
+
# beauty regressor). Falls back to the full image when InsightFace
|
| 269 |
+
# didn't find a face.
|
| 270 |
+
face_crop = _crop_to_face(img_array, insight_results.get("face_bbox"))
|
| 271 |
+
|
| 272 |
+
# Step 2: MediaPipe landmarks (works on the full image; it has its
|
| 273 |
+
# own internal detector).
|
| 274 |
+
logger.info("Running landmark analysis...")
|
| 275 |
+
landmark_results = landmarks.analyze(img_array)
|
| 276 |
+
results.update(landmark_results)
|
| 277 |
+
|
| 278 |
+
# Step 3: ethnicity classifier β likes a tighter face crop.
|
| 279 |
+
logger.info("Running ethnicity analysis...")
|
| 280 |
+
results.update(ethnicities.analyze(face_crop))
|
| 281 |
+
|
| 282 |
+
# Step 4: SegFormer parsing on the face crop (cleaner masks).
|
| 283 |
+
logger.info("Running face parsing...")
|
| 284 |
+
parse_results = parsing.analyze(face_crop)
|
| 285 |
+
results.update(parse_results)
|
| 286 |
+
|
| 287 |
+
# Step 5: HSEmotion on the face crop.
|
| 288 |
+
logger.info("Running emotion analysis...")
|
| 289 |
+
results.update(emotions.analyze(face_crop))
|
| 290 |
+
|
| 291 |
+
# Step 6: pixel-level colour analysis. Uses the face/hair masks
|
| 292 |
+
# from step 4 (already in face-crop coordinate space) and the
|
| 293 |
+
# MediaPipe lip/iris landmarks from step 2 (still in full-image
|
| 294 |
+
# space, normalised). We pass `face_crop` so mask coordinates
|
| 295 |
+
# line up; landmarks are in normalised coordinates so they map
|
| 296 |
+
# correctly to either image.
|
| 297 |
+
logger.info("Running color analysis...")
|
| 298 |
+
color_results = colors.analyze(
|
| 299 |
+
face_crop,
|
| 300 |
+
skin_mask=parse_results.get("_skin_mask"),
|
| 301 |
+
hair_mask=parse_results.get("_hair_mask"),
|
| 302 |
+
landmarks=landmark_results.get("_raw_landmarks"),
|
| 303 |
+
)
|
| 304 |
+
results.update(color_results)
|
| 305 |
+
|
| 306 |
+
# Step 7: obstruction classifier β also benefits from a face crop.
|
| 307 |
+
logger.info("Running obstruction analysis...")
|
| 308 |
+
results.update(obstructions.analyze(face_crop))
|
| 309 |
+
|
| 310 |
+
# Step 8: hair-type classifier.
|
| 311 |
+
logger.info("Running hair-type analysis...")
|
| 312 |
+
results.update(hair_types.analyze(face_crop))
|
| 313 |
+
|
| 314 |
+
# Step 9: learned beauty regressor (no-op if no weights present).
|
| 315 |
+
logger.info("Running beauty regressor...")
|
| 316 |
+
results.update(beauty.analyze(face_crop))
|
| 317 |
+
|
| 318 |
+
# Step 10: aesthetic aggregator. Reads the merged dict; no image
|
| 319 |
+
# input. Always runs last so it can see every other analyzer's
|
| 320 |
+
# outputs.
|
| 321 |
+
logger.info("Running aesthetic aggregator...")
|
| 322 |
+
results.update(aesthetics.analyze(results))
|
| 323 |
+
|
| 324 |
+
# Drop internal/scratch fields (leading underscore) before
|
| 325 |
+
# returning. Keeps masks and raw landmark lists out of the JSON.
|
| 326 |
+
results = {k: v for k, v in results.items() if not k.startswith("_")}
|
| 327 |
+
|
| 328 |
+
return results
|
| 329 |
+
|
| 330 |
+
|
| 331 |
@app.get("/")
|
| 332 |
async def root():
|
| 333 |
"""Service banner β confirms the server is reachable and which version."""
|
| 334 |
return {
|
| 335 |
"name": "HCP Face Analysis Service",
|
| 336 |
+
"version": "3.0.0",
|
| 337 |
"status": "running",
|
| 338 |
"endpoints": {
|
| 339 |
"health": "/health",
|
|
|
|
| 353 |
async def analyze_face(file: UploadFile = File(...)):
|
| 354 |
"""Multipart endpoint for direct uploads.
|
| 355 |
|
| 356 |
+
Runs the full ten-step pipeline and returns the merged attribute
|
| 357 |
+
dict. See `analyze_face_base64` for the JSON-body variant the
|
| 358 |
+
Express server calls.
|
| 359 |
"""
|
| 360 |
try:
|
|
|
|
|
|
|
|
|
|
| 361 |
contents = await file.read()
|
| 362 |
image = Image.open(io.BytesIO(contents)).convert("RGB")
|
| 363 |
img_array = np.array(image)
|
| 364 |
+
results = _run_pipeline(img_array)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 365 |
return {"success": True, "data": _to_json_safe(results)}
|
| 366 |
|
| 367 |
except Exception as e:
|
|
|
|
| 384 |
if not image_b64:
|
| 385 |
raise HTTPException(status_code=400, detail="No image data provided")
|
| 386 |
|
| 387 |
+
# Strip a possible "data:image/...;base64," prefix.
|
| 388 |
if "," in image_b64:
|
| 389 |
image_b64 = image_b64.split(",", 1)[1]
|
| 390 |
|
| 391 |
image_bytes = base64.b64decode(image_b64)
|
| 392 |
image = Image.open(io.BytesIO(image_bytes)).convert("RGB")
|
| 393 |
img_array = np.array(image)
|
| 394 |
+
results = _run_pipeline(img_array)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 395 |
return {"success": True, "data": _to_json_safe(results)}
|
| 396 |
|
| 397 |
except HTTPException:
|
architecture.md
CHANGED
|
@@ -2,65 +2,82 @@
|
|
| 2 |
|
| 3 |
## Pipeline
|
| 4 |
|
| 5 |
-
A single photo
|
| 6 |
-
into one dictionary; later analyzers overwrite
|
| 7 |
-
|
|
|
|
| 8 |
|
| 9 |
```
|
| 10 |
Photo (RGB ndarray)
|
| 11 |
β
|
| 12 |
-
βββΊ [1]
|
| 13 |
-
β
|
| 14 |
-
β
|
| 15 |
-
β smiling (mouthSmile blendshapes), eyes_open, possible_dimples,
|
| 16 |
-
β possible_unibrow, facial_asymmetry_score, blendshapes dict
|
| 17 |
β
|
| 18 |
-
βββΊ
|
| 19 |
-
β
|
| 20 |
-
β
|
| 21 |
β
|
| 22 |
-
βββΊ [
|
| 23 |
-
β
|
| 24 |
-
β
|
| 25 |
-
β
|
| 26 |
-
β
|
| 27 |
β
|
| 28 |
-
βββΊ [
|
| 29 |
-
β β
|
| 30 |
-
β
|
| 31 |
β
|
| 32 |
-
βββΊ [
|
| 33 |
-
β
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
β β skin_tone (Fitzpatrick + L*/a*/b* + hex), skin_undertone,
|
| 35 |
-
β eye_color, hair_color (name + hex), hair_texture
|
| 36 |
-
β lip_color (shade + hex)
|
| 37 |
β
|
| 38 |
-
βββΊ [
|
| 39 |
β β wearing_glasses, wearing_sunglasses, wearing_mask,
|
| 40 |
-
β
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
β
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
```
|
| 46 |
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
|
| 51 |
## Attribute β source map
|
| 52 |
|
| 53 |
-
The EditProfileScreen renders only fields backed by one of these
|
| 54 |
-
analyzers. Anything previously fed by the FaRL zero-shot classifier
|
| 55 |
-
has been removed because its outputs were too noisy to trust.
|
| 56 |
-
|
| 57 |
| Section | Field(s) | Source |
|
| 58 |
|---|---|---|
|
| 59 |
-
| Demographics |
|
| 60 |
-
|
|
| 61 |
-
|
|
| 62 |
-
|
|
| 63 |
-
| Hair |
|
|
|
|
| 64 |
| Hair | hair_color, hair hex | ColorAnalyzer |
|
| 65 |
| Eyes | eye_shape, eye_depth, eye_spacing, eye_size, eyes_open | MediaPipe |
|
| 66 |
| Eyes | eye_color | ColorAnalyzer |
|
|
@@ -70,30 +87,58 @@ has been removed because its outputs were too noisy to trust.
|
|
| 70 |
| Lips & Mouth | lip_color (shade + hex) | ColorAnalyzer (mask from MediaPipe) |
|
| 71 |
| Skin | skin_tone (Fitzpatrick, L*/a*/b*, hex), skin_undertone | ColorAnalyzer |
|
| 72 |
| Skin | wrinkle_level, skin_texture_score, skin_uniformity, freckles_or_moles | SegFormer mask + OpenCV stats |
|
| 73 |
-
| Accessories | wearing_glasses, wearing_sunglasses, wearing_mask | ObstructionViT |
|
| 74 |
| Accessories | wearing_hat | SegFormer (hat class coverage) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
## Deployment
|
| 77 |
|
| 78 |
-
The service
|
| 79 |
-
free tier (
|
| 80 |
-
|
| 81 |
-
|
|
|
|
| 82 |
|
| 83 |
The Node/Express server forwards `/analyze-face` requests to
|
| 84 |
-
`FACE_SERVICE_URL/analyze-base64`. The React Native client never
|
| 85 |
-
to this service directly.
|
| 86 |
|
| 87 |
## Adding a new analyzer
|
| 88 |
|
| 89 |
-
1. Drop a new module under `analyzers/`
|
| 90 |
-
`__init__()` and `analyze(
|
| 91 |
-
2. Import
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
|
|
|
| 96 |
|
| 97 |
Order matters: later analyzers overwrite earlier keys on collision.
|
| 98 |
-
The
|
| 99 |
-
signal.
|
|
|
|
| 2 |
|
| 3 |
## Pipeline
|
| 4 |
|
| 5 |
+
A single photo runs through ten analyzers. Their outputs are merged
|
| 6 |
+
into one dictionary; later analyzers can overwrite keys from earlier
|
| 7 |
+
ones (only intentional in a couple of places β `_run_pipeline` in
|
| 8 |
+
[app.py](app.py) is the single source of truth).
|
| 9 |
|
| 10 |
```
|
| 11 |
Photo (RGB ndarray)
|
| 12 |
β
|
| 13 |
+
βββΊ [1] InsightFaceAnalyzer (insightface buffalo_l, ONNX)
|
| 14 |
+
β β face_bbox, face_confidence, face_embedding (512-d ArcFace),
|
| 15 |
+
β age_estimate, age_range, gender + confidences
|
|
|
|
|
|
|
| 16 |
β
|
| 17 |
+
βββΊ Build face crop from face_bbox + padding. Downstream analyzers
|
| 18 |
+
β that benefit from a tighter input read the crop; MediaPipe gets
|
| 19 |
+
β the full image because it has its own detector.
|
| 20 |
β
|
| 21 |
+
βββΊ [2] LandmarkAnalyzer (MediaPipe Face Landmarker)
|
| 22 |
+
β 478 landmarks + 52 blendshapes β all geometric features,
|
| 23 |
+
β smiling, mouth_open (via blendshapes.jawOpen), eyes_open,
|
| 24 |
+
β facial_asymmetry_score, smile_asymmetry, possible_dimples,
|
| 25 |
+
β possible_unibrow.
|
| 26 |
β
|
| 27 |
+
βββΊ [3] EthnicityAnalyzer (cledoux42/Ethnicity_Test_v003 ViT)
|
| 28 |
+
β β ethnicity, ethnicity_confidence, ethnicity_distribution
|
| 29 |
+
β (cropped input).
|
| 30 |
β
|
| 31 |
+
βββΊ [4] ParsingAnalyzer (SegFormer-B5 human parsing)
|
| 32 |
+
β β _skin_mask, _hair_mask, hat_detected, hair_length,
|
| 33 |
+
β hair_present, wrinkle_level, skin_texture_score,
|
| 34 |
+
β skin_uniformity, freckles_or_moles
|
| 35 |
+
β (cropped input β cleaner masks).
|
| 36 |
+
β
|
| 37 |
+
βββΊ [5] EmotionAnalyzer (HSEmotion EfficientNet-B0)
|
| 38 |
+
β β primary/secondary emotion, emotion_scores, valence,
|
| 39 |
+
β arousal, mood (cropped input).
|
| 40 |
+
β
|
| 41 |
+
βββΊ [6] ColorAnalyzer (no ML β OpenCV LAB/HSV)
|
| 42 |
+
β Reads SegFormer masks + MediaPipe lip/iris landmarks.
|
| 43 |
β β skin_tone (Fitzpatrick + L*/a*/b* + hex), skin_undertone,
|
| 44 |
+
β eye_color, hair_color (name + hex), hair_texture
|
| 45 |
+
β (coarse, fallback), lip_color (shade + hex)
|
| 46 |
β
|
| 47 |
+
βββΊ [7] ObstructionAnalyzer (dima806/face_obstruction ViT-B/16)
|
| 48 |
β β wearing_glasses, wearing_sunglasses, wearing_mask,
|
| 49 |
+
β obstruction_scores (cropped input).
|
| 50 |
+
β
|
| 51 |
+
βββΊ [8] HairTypeAnalyzer (dima806/hair_type ViT-B/16)
|
| 52 |
+
β β hair_type (curly/dreadlocks/kinky/straight/wavy),
|
| 53 |
+
β hair_type_confidence (cropped input).
|
| 54 |
β
|
| 55 |
+
βββΊ [9] BeautyAnalyzer (ResNet-50 trained on SCUT-FBP5500)
|
| 56 |
+
β Optional. Loads local weights or HF Hub; if absent, output
|
| 57 |
+
β is None and AestheticAnalyzer falls back to rules.
|
| 58 |
+
β β beauty_score (1.0β5.0), beauty_score_norm (0β100),
|
| 59 |
+
β beauty_model_source.
|
| 60 |
+
β
|
| 61 |
+
βββΊ [10] AestheticAnalyzer (no model)
|
| 62 |
+
Reads the merged dict from steps 1β9 and produces the
|
| 63 |
+
final chopped_score (0β100) plus chopped_breakdown
|
| 64 |
+
showing each factor's signed contribution.
|
| 65 |
```
|
| 66 |
|
| 67 |
+
Internal/scratch keys use a leading underscore (`_skin_mask`,
|
| 68 |
+
`_hair_mask`, `_raw_landmarks`, `_insight_landmarks_2d`). `app.py`
|
| 69 |
+
strips them before returning JSON.
|
| 70 |
|
| 71 |
## Attribute β source map
|
| 72 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
| Section | Field(s) | Source |
|
| 74 |
|---|---|---|
|
| 75 |
+
| Demographics | face_bbox, face_confidence, face_embedding (512-d), age_estimate, age_range, age_confidence, gender, gender_confidence | InsightFace buffalo_l |
|
| 76 |
+
| Demographics | ethnicity, ethnicity_confidence, ethnicity_distribution | EthnicityAnalyzer (cledoux42 ViT) |
|
| 77 |
+
| Emotion | primary/secondary emotion, emotion_scores, valence, arousal, mood | HSEmotion EffNet-B0 |
|
| 78 |
+
| Face Structure | face_shape (+ 4 ratios), jawline_type/angle, chin_type, cheekbone_prominence, cheek_fullness, forehead_width, facial_asymmetry_score | MediaPipe Face Landmarker |
|
| 79 |
+
| Hair | hair_length, hair_present | SegFormer-B5 |
|
| 80 |
+
| Hair | hair_type (+ confidence) | HairTypeViT (dima806) |
|
| 81 |
| Hair | hair_color, hair hex | ColorAnalyzer |
|
| 82 |
| Eyes | eye_shape, eye_depth, eye_spacing, eye_size, eyes_open | MediaPipe |
|
| 83 |
| Eyes | eye_color | ColorAnalyzer |
|
|
|
|
| 87 |
| Lips & Mouth | lip_color (shade + hex) | ColorAnalyzer (mask from MediaPipe) |
|
| 88 |
| Skin | skin_tone (Fitzpatrick, L*/a*/b*, hex), skin_undertone | ColorAnalyzer |
|
| 89 |
| Skin | wrinkle_level, skin_texture_score, skin_uniformity, freckles_or_moles | SegFormer mask + OpenCV stats |
|
| 90 |
+
| Accessories | wearing_glasses, wearing_sunglasses, wearing_mask | ObstructionViT (dima806) |
|
| 91 |
| Accessories | wearing_hat | SegFormer (hat class coverage) |
|
| 92 |
+
| Aesthetics | beauty_score (1β5), beauty_score_norm (0β100) | BeautyAnalyzer (SCUT-FBP5500 ResNet-50) |
|
| 93 |
+
| Aesthetics | chopped_score (0β100), chopped_breakdown | AestheticAnalyzer (rule + learned blend) |
|
| 94 |
+
|
| 95 |
+
## Face matching
|
| 96 |
+
|
| 97 |
+
InsightFace's ArcFace head emits a 512-d L2-normalised recognition
|
| 98 |
+
embedding. We store it alongside each contact in
|
| 99 |
+
`people.face_embedding` (pgvector). On a new photo save, the client
|
| 100 |
+
queries Supabase for any contact with cosine similarity β₯ 0.55 to the
|
| 101 |
+
new embedding and prompts the user *"this looks like {name}, add to
|
| 102 |
+
that profile?"* before creating a new contact.
|
| 103 |
+
|
| 104 |
+
LFW accuracy is 99.83%; IJB-B at FAR=1e-4 is 96.21%. For grouping
|
| 105 |
+
photos in a personal collection (similar lighting, same camera) this
|
| 106 |
+
is excellent. Identical twins and close family members can match β the
|
| 107 |
+
0.55 threshold makes the prompt opt-in rather than auto-merge.
|
| 108 |
+
|
| 109 |
+
## Training the beauty regressor
|
| 110 |
+
|
| 111 |
+
Live source in [training/beauty/](../training/beauty/). The script
|
| 112 |
+
fine-tunes a timm ResNet-50 on SCUT-FBP5500. After training, drop the
|
| 113 |
+
resulting `beauty_regressor.pt` into `face-service/models/` (or push
|
| 114 |
+
to HF Hub and set `BEAUTY_HF_REPO_ID`). `BeautyAnalyzer` picks it up
|
| 115 |
+
automatically on the next process boot.
|
| 116 |
+
|
| 117 |
+
Until weights exist, `beauty_score` returns None and the AestheticAnalyzer
|
| 118 |
+
gracefully falls back to a pure rule-based chopped score.
|
| 119 |
|
| 120 |
## Deployment
|
| 121 |
|
| 122 |
+
The service builds as a Docker image targeting Hugging Face Spaces
|
| 123 |
+
free tier (2 GB RAM, shared CPU). MediaPipe `.task` and the
|
| 124 |
+
InsightFace buffalo_l bundle are pulled at build time; all other
|
| 125 |
+
Hugging Face models lazy-download on first inference and cache under
|
| 126 |
+
`/root/.cache/huggingface`.
|
| 127 |
|
| 128 |
The Node/Express server forwards `/analyze-face` requests to
|
| 129 |
+
`FACE_SERVICE_URL/analyze-base64`. The React Native client never
|
| 130 |
+
talks to this service directly.
|
| 131 |
|
| 132 |
## Adding a new analyzer
|
| 133 |
|
| 134 |
+
1. Drop a new module under `analyzers/` with a class exposing
|
| 135 |
+
`__init__()` and `analyze(...) -> dict`.
|
| 136 |
+
2. Import + add a lazy-load block in `app.py`'s `get_analyzers()`.
|
| 137 |
+
3. Add a `results.update(...)` call inside `_run_pipeline` at the
|
| 138 |
+
right pipeline position.
|
| 139 |
+
4. Surface the new keys in
|
| 140 |
+
[client/src/screens/EditProfileScreen.js](../client/src/screens/EditProfileScreen.js)
|
| 141 |
+
and add a legend row.
|
| 142 |
|
| 143 |
Order matters: later analyzers overwrite earlier keys on collision.
|
| 144 |
+
The aesthetic aggregator runs last so it can see everything.
|
|
|
requirements.txt
CHANGED
|
@@ -13,3 +13,5 @@ timm==1.0.3
|
|
| 13 |
safetensors>=0.6.0
|
| 14 |
transformers==4.45.2
|
| 15 |
hsemotion>=0.2.2
|
|
|
|
|
|
|
|
|
| 13 |
safetensors>=0.6.0
|
| 14 |
transformers==4.45.2
|
| 15 |
hsemotion>=0.2.2
|
| 16 |
+
insightface>=0.7.3
|
| 17 |
+
onnxruntime>=1.18.0
|