Evan Li commited on
Commit
8f19f34
Β·
1 Parent(s): ee3a08a
Dockerfile CHANGED
@@ -2,7 +2,7 @@ FROM python:3.11-slim
2
 
3
  # System deps for OpenCV and MediaPipe
4
  RUN apt-get update && apt-get install -y --no-install-recommends \
5
- libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 wget \
6
  libegl1 libgles2 libgomp1 \
7
  g++ && \
8
  rm -rf /var/lib/apt/lists/*
@@ -14,13 +14,25 @@ COPY requirements.txt .
14
  RUN pip install --no-cache-dir -r requirements.txt
15
 
16
  # Pre-download MediaPipe model at build time so first request is fast.
17
- # All other models (FairFace, SegFormer, HSEmotion, ObstructionViT,
18
- # HairTypeViT) are pulled from Hugging Face on first request and cached
19
- # in /root/.cache/huggingface for the rest of the process lifetime.
 
20
  RUN mkdir -p models && \
21
  wget -q -O models/face_landmarker.task \
22
  "https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/latest/face_landmarker.task"
23
 
 
 
 
 
 
 
 
 
 
 
 
24
  COPY . .
25
 
26
  EXPOSE 7860
 
2
 
3
  # System deps for OpenCV and MediaPipe
4
  RUN apt-get update && apt-get install -y --no-install-recommends \
5
+ libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 wget unzip \
6
  libegl1 libgles2 libgomp1 \
7
  g++ && \
8
  rm -rf /var/lib/apt/lists/*
 
14
  RUN pip install --no-cache-dir -r requirements.txt
15
 
16
  # Pre-download MediaPipe model at build time so first request is fast.
17
+ # Hugging Face models (Ethnicity ViT, SegFormer, HSEmotion, ObstructionViT,
18
+ # HairTypeViT) and the InsightFace buffalo_l bundle are pulled lazily on
19
+ # first request and cached in /root/.cache for the lifetime of the
20
+ # container.
21
  RUN mkdir -p models && \
22
  wget -q -O models/face_landmarker.task \
23
  "https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/latest/face_landmarker.task"
24
 
25
+ # Pre-download InsightFace buffalo_l bundle (detection + recognition +
26
+ # age + gender + landmarks) so the first /analyze call doesn't pay the
27
+ # ~280MB download. The bundle auto-extracts under ~/.insightface/models/
28
+ # on first use.
29
+ RUN mkdir -p /root/.insightface/models && \
30
+ wget -q -O /root/.insightface/models/buffalo_l.zip \
31
+ "https://github.com/deepinsight/insightface/releases/download/v0.7/buffalo_l.zip" && \
32
+ cd /root/.insightface/models && unzip -q buffalo_l.zip -d buffalo_l && rm buffalo_l.zip
33
+
34
+ # unzip wasn't in the system deps; add it via the apt block at the top.
35
+
36
  COPY . .
37
 
38
  EXPOSE 7860
README.md CHANGED
@@ -10,26 +10,33 @@ pinned: false
10
 
11
  # HCP Face Analysis Microservice
12
 
13
- FastAPI service that runs seven specialized analyzers over a single photo
14
- and returns a merged dictionary of ~100 facial attributes.
 
15
 
16
  ## Models
17
 
18
  | # | Component | Model | Task | Size |
19
  |---|-----------|-------|------|------|
20
- | 1 | MediaPipe Face Landmarker | `face_landmarker.task` (Google) | 478 3D landmarks + 52 ARKit blendshapes β€” geometric features, smiling, mouth-open | ~4 MB |
21
- | 2 | FairFace age | `dima806/fairface_age_image_detection` (ViT-B/16) | 9-bucket age β†’ softmax-weighted continuous estimate | ~340 MB |
22
- | 2 | FairFace gender | `dima806/fairface_gender_image_detection` (ViT-B/16) | Binary gender (~93.4% acc) | ~340 MB |
23
- | 2 | Ethnicity | `cledoux42/Ethnicity_Test_v003` (ViT) | 5-class ethnicity (~79.6% acc) | ~340 MB |
24
- | 3 | Human parsing | `matei-dorian/segformer-b5-finetuned-human-parsing` | 18-class pixel segmentation β†’ masks + hair length + hat | ~340 MB |
25
- | 4 | Emotion | HSEmotion `enet_b0_8_best_afew` (EfficientNet-B0) | 8-class emotion + valence/arousal | ~20 MB |
26
- | 5 | Color analysis | (no model β€” OpenCV LAB/HSV) | Skin tone, hair color, eye color, lip color | 0 MB |
27
- | 6 | Obstruction | `dima806/face_obstruction_image_detection` (ViT-B/16) | glasses / sunglasses / mask (~99% precision) | ~340 MB |
28
- | 7 | Hair type | `dima806/hair_type_image_detection` (ViT-B/16) | curly/dreadlocks/kinky/straight/wavy (~93% acc) | ~340 MB |
29
-
30
- All analyzers are lazy-loaded on first request. The MediaPipe weight
31
- file is pre-downloaded at Docker build time; all Hugging Face models
32
- are cached on first inference.
 
 
 
 
 
 
33
 
34
  ## API endpoints
35
 
@@ -46,5 +53,5 @@ curl -X POST https://YOUR-SPACE.hf.space/analyze-base64 \
46
  -d '{"image": "<base64-encoded-image>"}'
47
  ```
48
 
49
- See [architecture.md](./architecture.md) for the pipeline diagram and the
50
- full per-attribute model attribution table.
 
10
 
11
  # HCP Face Analysis Microservice
12
 
13
+ FastAPI service that runs ten specialised analyzers over a single
14
+ photo and returns a merged dictionary of facial attributes plus a
15
+ face-recognition embedding and an aesthetic "chopped score."
16
 
17
  ## Models
18
 
19
  | # | Component | Model | Task | Size |
20
  |---|-----------|-------|------|------|
21
+ | 1 | InsightFace | `buffalo_l` (SCRFD + ArcFace ResNet50, ONNX) | Detection + 512-d recognition embedding + age + gender + 106 landmarks (99.83% LFW) | ~280 MB |
22
+ | 2 | MediaPipe Face Landmarker | `face_landmarker.task` (Google) | 478 3D landmarks + 52 ARKit blendshapes β€” geometric features, smiling, mouth-open | ~4 MB |
23
+ | 3 | Ethnicity | `cledoux42/Ethnicity_Test_v003` (ViT) | 5-class ethnicity (~79.6% acc) | ~340 MB |
24
+ | 4 | Human parsing | `matei-dorian/segformer-b5-finetuned-human-parsing` | 18-class pixel segmentation β†’ masks + hair length + hat | ~340 MB |
25
+ | 5 | Emotion | HSEmotion `enet_b0_8_best_afew` (EfficientNet-B0) | 8-class emotion + valence/arousal | ~20 MB |
26
+ | 6 | Color analysis | (no model β€” OpenCV LAB/HSV) | Skin tone, hair color, eye color, lip color | 0 MB |
27
+ | 7 | Obstruction | `dima806/face_obstruction_image_detection` (ViT-B/16) | glasses / sunglasses / mask (~99% precision) | ~340 MB |
28
+ | 8 | Hair type | `dima806/hair_type_image_detection` (ViT-B/16) | curly/dreadlocks/kinky/straight/wavy (~93% acc) | ~340 MB |
29
+ | 9 | Beauty regression | timm ResNet-50, fine-tuned on SCUT-FBP5500 | 1.0–5.0 beauty score (~Pearson r β‰₯ 0.85 expected) | ~100 MB |
30
+ | 10 | Aesthetic aggregator | (no model β€” Python rules) | Combines learned beauty + heuristic factors β†’ chopped_score (0–100) | 0 MB |
31
+
32
+ InsightFace `buffalo_l` is pre-downloaded at Docker build time. The
33
+ MediaPipe weight file is pre-downloaded too. All Hugging Face models
34
+ lazy-load on first inference and cache locally for the lifetime of
35
+ the process. The SCUT-FBP5500-trained beauty regressor must be
36
+ produced via [`training/beauty/`](../training/beauty/) β€” until you
37
+ drop weights in `models/beauty_regressor.pt` (or set
38
+ `BEAUTY_HF_REPO_ID`), `beauty_score` returns null and the chopped
39
+ score uses heuristics only.
40
 
41
  ## API endpoints
42
 
 
53
  -d '{"image": "<base64-encoded-image>"}'
54
  ```
55
 
56
+ See [architecture.md](./architecture.md) for the pipeline diagram and
57
+ the full per-attribute β†’ source map.
analyzers/aesthetic_analyzer.py ADDED
@@ -0,0 +1,220 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ AestheticAnalyzer β€” "chopped score" aggregator.
3
+
4
+ What it does
5
+ ------------
6
+ Reads the merged result dict from every other analyzer and produces a
7
+ single numeric "chopped score" plus a per-factor breakdown. Higher
8
+ score = more chopped = less conventionally attractive (by the
9
+ arbitrary rubric encoded here). The breakdown lets you tune weights
10
+ or flip polarity client-side without rerunning inference.
11
+
12
+ Score composition
13
+ -----------------
14
+ The final chopped_score is a weighted blend of two sources:
15
+
16
+ 1. **Learned beauty regressor** (from BeautyAnalyzer, trained on
17
+ SCUT-FBP5500): a number in [1.0, 5.0] reflecting averaged human
18
+ ratings. We rescale to a 0–100 unattractiveness axis. This is the
19
+ dominant signal when available β€” heavy weight (default 0.7).
20
+
21
+ 2. **Rule-based factor sum**: penalties for asymmetry, wrinkles,
22
+ uneven skin, freckles, and asymmetric smile; bonuses for defined
23
+ jawline, prominent cheekbones, clear skin, balanced lips, and
24
+ dimples. Each factor is documented in `_compute_rule_score`.
25
+ This is the only signal when the regressor isn't loaded
26
+ (BeautyAnalyzer returns None).
27
+
28
+ Blend math
29
+ ----------
30
+ if beauty_score available:
31
+ chopped = 0.7 * (100 - beauty_norm) + 0.3 * rule_score
32
+ else:
33
+ chopped = rule_score
34
+ chopped is clamped to [0, 100].
35
+
36
+ Subjectivity disclaimer
37
+ -----------------------
38
+ Every weight in this file is a guess. "Beauty" is subjective, culturally
39
+ biased, and reductive. Treat the score as an in-joke metric; never
40
+ expose it as objective truth. The UI gates the row behind a
41
+ Settings toggle off-by-default for that reason.
42
+
43
+ Note: this analyzer takes no image input β€” it reads the merged result
44
+ dict produced by every other analyzer that ran ahead of it.
45
+ """
46
+
47
+ from typing import Any
48
+
49
+
50
+ # How much weight the learned beauty regressor gets when both signals
51
+ # are available. The rule-based sum gets the rest (1 - this).
52
+ LEARNED_WEIGHT = 0.7
53
+
54
+ # Baseline score. Penalties push up, bonuses pull down.
55
+ BASELINE = 50.0
56
+
57
+
58
+ class AestheticAnalyzer:
59
+ def __init__(self):
60
+ # No model to load.
61
+ pass
62
+
63
+ def analyze(self, merged: dict[str, Any]) -> dict[str, Any]:
64
+ """Compute chopped score from the merged result dict.
65
+
66
+ Unusual signature: not (img_rgb), since this analyzer aggregates
67
+ prior results rather than running inference on the image.
68
+ app.py special-cases the call to pass `merged` here.
69
+ """
70
+ rule_score, breakdown = self._compute_rule_score(merged)
71
+
72
+ beauty_norm = merged.get("beauty_score_norm")
73
+ if beauty_norm is not None:
74
+ # Beauty regressor: 0 = ugly, 100 = beautiful (per SCUT-FBP5500
75
+ # scaling). Flip to unattractiveness axis: 100 - x.
76
+ learned_unattractive = 100.0 - float(beauty_norm)
77
+ chopped = (
78
+ LEARNED_WEIGHT * learned_unattractive
79
+ + (1.0 - LEARNED_WEIGHT) * rule_score
80
+ )
81
+ breakdown["learned_unattractive"] = round(
82
+ LEARNED_WEIGHT * learned_unattractive - LEARNED_WEIGHT * BASELINE, 2
83
+ )
84
+ breakdown["_blend_weight_learned"] = LEARNED_WEIGHT
85
+ else:
86
+ chopped = rule_score
87
+ breakdown["_blend_weight_learned"] = 0.0
88
+
89
+ chopped = max(0.0, min(100.0, chopped))
90
+
91
+ return {
92
+ "chopped_score": round(chopped, 1),
93
+ "chopped_breakdown": breakdown,
94
+ "chopped_polarity_note": (
95
+ "0 = least chopped, 100 = most chopped. "
96
+ "Subtract from 100 for an 'attractiveness' read."
97
+ ),
98
+ }
99
+
100
+ # ------------------------------------------------------------------
101
+ # Rule-based scoring
102
+ # ------------------------------------------------------------------
103
+
104
+ @staticmethod
105
+ def _compute_rule_score(d: dict[str, Any]) -> tuple[float, dict[str, float]]:
106
+ """Hand-tuned weighted sum over previously-extracted attributes.
107
+
108
+ Returns (score, breakdown_dict). The breakdown gives each factor's
109
+ signed contribution so a UI can show *why* a score landed where
110
+ it did. Score starts at BASELINE (50) and moves up/down.
111
+ """
112
+ score = BASELINE
113
+ breakdown: dict[str, float] = {}
114
+
115
+ # ── Penalties (push score up = more chopped) ─────────────────
116
+
117
+ # Facial asymmetry: 0 = perfectly symmetric, 1 = very asymmetric.
118
+ # MediaPipe `facial_asymmetry_score` is already in this range.
119
+ asym = d.get("facial_asymmetry_score")
120
+ if isinstance(asym, (int, float)):
121
+ penalty = float(asym) * 18.0
122
+ score += penalty
123
+ breakdown["asymmetry_penalty"] = round(penalty, 2)
124
+
125
+ # Wrinkle level from SegFormer + OpenCV Laplacian classification.
126
+ wrinkle_penalty_map = {
127
+ "smooth": 0.0, "slight": 4.0, "moderate": 8.0, "prominent": 12.0,
128
+ }
129
+ wrinkle = d.get("wrinkle_level")
130
+ if wrinkle in wrinkle_penalty_map:
131
+ penalty = wrinkle_penalty_map[wrinkle]
132
+ score += penalty
133
+ breakdown["wrinkle_penalty"] = penalty
134
+
135
+ # Skin uniformity = LAB L* std-dev over the face mask. Higher
136
+ # std means uneven tone (shadows, blemishes). Scale up to +8.
137
+ uniformity = d.get("skin_uniformity")
138
+ if isinstance(uniformity, (int, float)) and uniformity > 0:
139
+ # Empirically, uniformity in clean skin is ~8-15; very uneven
140
+ # skin pushes into the 20-30 range.
141
+ penalty = min(8.0, max(0.0, (float(uniformity) - 10.0) * 0.5))
142
+ score += penalty
143
+ breakdown["skin_unevenness_penalty"] = round(penalty, 2)
144
+
145
+ # Freckles/moles bucket.
146
+ freckle_penalty_map = {"none": 0.0, "few": 1.0, "some": 3.0, "many": 5.0}
147
+ freckles = d.get("freckles_or_moles")
148
+ if freckles in freckle_penalty_map:
149
+ penalty = freckle_penalty_map[freckles]
150
+ score += penalty
151
+ breakdown["freckles_penalty"] = penalty
152
+
153
+ # Smile asymmetry: 0 = perfectly symmetric smile, larger = lopsided.
154
+ smile_asym = d.get("smile_asymmetry")
155
+ if isinstance(smile_asym, (int, float)):
156
+ penalty = min(6.0, float(smile_asym) * 30.0)
157
+ score += penalty
158
+ breakdown["smile_asymmetry_penalty"] = round(penalty, 2)
159
+
160
+ # Photo-quality penalty: sunglasses/mask hide features and the
161
+ # model is guessing more. Mild penalty, not a personal trait.
162
+ if d.get("wearing_sunglasses") or d.get("wearing_mask"):
163
+ score += 5.0
164
+ breakdown["obstruction_penalty"] = 5.0
165
+
166
+ # ── Bonuses (pull score down = less chopped) ─────────────────
167
+
168
+ # Defined jawline. Two signals (string bucket + numeric angle);
169
+ # take the stronger of the two contributions.
170
+ jaw_bonus = 0.0
171
+ jaw_type = d.get("jawline_type")
172
+ jaw_type_bonus_map = {"sharp": -10.0, "strong": -6.0, "soft": 0.0}
173
+ if jaw_type in jaw_type_bonus_map:
174
+ jaw_bonus = jaw_type_bonus_map[jaw_type]
175
+ jaw_angle = d.get("jawline_angle")
176
+ if isinstance(jaw_angle, (int, float)) and jaw_angle < 115:
177
+ # Sharp angles add more on top of the categorical signal.
178
+ jaw_bonus = min(jaw_bonus, -10.0)
179
+ if jaw_bonus:
180
+ score += jaw_bonus
181
+ breakdown["jaw_definition_bonus"] = round(jaw_bonus, 2)
182
+
183
+ # Cheekbone prominence.
184
+ cheek_bonus_map = {"high": -7.0, "moderate": -3.0, "flat": 0.0}
185
+ cheek = d.get("cheekbone_prominence")
186
+ if cheek in cheek_bonus_map:
187
+ bonus = cheek_bonus_map[cheek]
188
+ score += bonus
189
+ breakdown["cheekbone_bonus"] = bonus
190
+
191
+ # Skin clarity bonus when the texture score is low (i.e. smooth skin).
192
+ # skin_texture_score is the same Laplacian-density value used by
193
+ # wrinkle_level; ≀4 is "smooth" territory.
194
+ texture = d.get("skin_texture_score")
195
+ if isinstance(texture, (int, float)) and 0 < texture <= 4:
196
+ score -= 9.0
197
+ breakdown["skin_clarity_bonus"] = -9.0
198
+
199
+ # Lip fullness β€” "average" and "full" both read as healthy.
200
+ lip = d.get("lip_fullness")
201
+ if lip in {"average", "full"}:
202
+ score -= 5.0
203
+ breakdown["lip_fullness_bonus"] = -5.0
204
+
205
+ # Defined cupid's bow.
206
+ if d.get("cupids_bow") == "defined":
207
+ score -= 3.0
208
+ breakdown["cupids_bow_bonus"] = -3.0
209
+
210
+ # Normal eye spacing.
211
+ if d.get("eye_spacing") == "average":
212
+ score -= 4.0
213
+ breakdown["eye_spacing_bonus"] = -4.0
214
+
215
+ # Dimples β€” small bonus when the MediaPipe heuristic fires.
216
+ if d.get("possible_dimples"):
217
+ score -= 3.0
218
+ breakdown["dimples_bonus"] = -3.0
219
+
220
+ return score, breakdown
analyzers/beauty_analyzer.py ADDED
@@ -0,0 +1,172 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ BeautyAnalyzer β€” learned facial-beauty regression on SCUT-FBP5500.
3
+
4
+ Model
5
+ -----
6
+ - Architecture : timm ResNet-50 with a single-output regression head
7
+ (output value in the SCUT-FBP5500 1.0–5.0 score range,
8
+ averaged from 60 human raters per image).
9
+ - Dataset : SCUT-FBP5500 (5,500 faces, gender + race balanced).
10
+ https://github.com/HCIILAB/SCUT-FBP5500-Database-Release
11
+ - Trained by : the user, via the training kit at `training/beauty/`.
12
+ - Expected MAE : ~0.25–0.30 on the standard test split; Pearson r β‰₯ 0.85.
13
+ - License : training data is research-only; the trained weights
14
+ you produce are yours to license as you choose.
15
+
16
+ Weight loading
17
+ --------------
18
+ Two ways the analyzer finds weights, tried in order:
19
+ 1. Local file at `models/beauty_regressor.pt` (drop in after training).
20
+ 2. Hugging Face Hub repo, controlled by `BEAUTY_HF_REPO_ID` env var
21
+ (e.g. `your-username/scut-fbp5500-resnet50`), loaded via
22
+ huggingface_hub.hf_hub_download.
23
+
24
+ If neither resolves, the analyzer logs a warning and returns
25
+ `beauty_score: None`, which AestheticAnalyzer detects and falls back
26
+ to the pure rule-based chopped score.
27
+
28
+ Inputs
29
+ ------
30
+ img_rgb : np.ndarray (H, W, 3) uint8
31
+
32
+ Outputs (dict)
33
+ --------------
34
+ beauty_score : float in [1.0, 5.0] (SCUT-FBP5500 native range)
35
+ or None if no model is available
36
+ beauty_score_norm : float in [0.0, 100.0] (linearly rescaled)
37
+ beauty_model_source : "local" | "huggingface" | "unavailable"
38
+ """
39
+
40
+ import os
41
+ from typing import Any
42
+
43
+ import numpy as np
44
+ from PIL import Image
45
+
46
+ try:
47
+ import torch
48
+ import timm
49
+ from torchvision import transforms
50
+ HAS_TORCH = True
51
+ except ImportError:
52
+ HAS_TORCH = False
53
+
54
+
55
+ LOCAL_WEIGHTS_PATH = os.environ.get(
56
+ "BEAUTY_WEIGHTS_PATH", "models/beauty_regressor.pt"
57
+ )
58
+ HF_REPO_ID = os.environ.get("BEAUTY_HF_REPO_ID") # e.g. "user/scut-fbp5500-resnet50"
59
+ HF_FILENAME = os.environ.get("BEAUTY_HF_FILENAME", "beauty_regressor.pt")
60
+ BACKBONE = os.environ.get("BEAUTY_BACKBONE", "resnet50")
61
+
62
+ # Standard ImageNet stats β€” SCUT-FBP5500 fine-tunes from ImageNet-pretrained
63
+ # backbones so we use the same normalisation at inference time.
64
+ IMAGENET_MEAN = [0.485, 0.456, 0.406]
65
+ IMAGENET_STD = [0.229, 0.224, 0.225]
66
+
67
+
68
+ class BeautyAnalyzer:
69
+ def __init__(self):
70
+ self.device = (
71
+ torch.device("cuda" if torch.cuda.is_available() else "cpu")
72
+ if HAS_TORCH else None
73
+ )
74
+ self.model = None
75
+ self.source = "unavailable"
76
+ self.transform = None
77
+
78
+ if not HAS_TORCH:
79
+ print(
80
+ "[BeautyAnalyzer] torch / timm not installed β€” beauty_score "
81
+ "will be None and AestheticAnalyzer will fall back to rules."
82
+ )
83
+ return
84
+
85
+ weights_path = self._resolve_weights_path()
86
+ if weights_path is None:
87
+ print(
88
+ "[BeautyAnalyzer] No trained weights found at "
89
+ f"{LOCAL_WEIGHTS_PATH} and BEAUTY_HF_REPO_ID is unset. "
90
+ "Train one via `training/beauty/train.py` and drop the "
91
+ ".pt into face-service/models/, or set BEAUTY_HF_REPO_ID."
92
+ )
93
+ return
94
+
95
+ try:
96
+ # Build backbone with a single regression output. Matches the
97
+ # training script's architecture exactly β€” see
98
+ # training/beauty/train.py.
99
+ self.model = timm.create_model(BACKBONE, pretrained=False, num_classes=1)
100
+ state = torch.load(weights_path, map_location=self.device)
101
+ # Support both bare state_dicts and {"state_dict": ...} wrappers.
102
+ if isinstance(state, dict) and "state_dict" in state:
103
+ state = state["state_dict"]
104
+ self.model.load_state_dict(state, strict=True)
105
+ self.model.to(self.device).eval()
106
+
107
+ # SCUT-FBP5500 standard inference transform: 224Γ—224, ImageNet norm.
108
+ self.transform = transforms.Compose([
109
+ transforms.Resize((224, 224)),
110
+ transforms.ToTensor(),
111
+ transforms.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD),
112
+ ])
113
+ print(
114
+ f"[BeautyAnalyzer] Loaded {BACKBONE} weights from "
115
+ f"{weights_path} (source={self.source})"
116
+ )
117
+ except Exception as exc:
118
+ print(f"[BeautyAnalyzer] Failed to load beauty model: {exc}")
119
+ self.model = None
120
+
121
+ @staticmethod
122
+ def _resolve_weights_path() -> str | None:
123
+ """Local file wins, HF Hub is the fallback."""
124
+ if os.path.exists(LOCAL_WEIGHTS_PATH):
125
+ return LOCAL_WEIGHTS_PATH
126
+ if HF_REPO_ID:
127
+ try:
128
+ from huggingface_hub import hf_hub_download
129
+ return hf_hub_download(repo_id=HF_REPO_ID, filename=HF_FILENAME)
130
+ except Exception as exc:
131
+ print(f"[BeautyAnalyzer] HF Hub download failed: {exc}")
132
+ return None
133
+
134
+ def _set_source(self, source: str) -> None:
135
+ self.source = source
136
+
137
+ def analyze(self, img_rgb: np.ndarray) -> dict[str, Any]:
138
+ if self.model is None or self.transform is None:
139
+ return self._empty_result()
140
+
141
+ try:
142
+ pil = Image.fromarray(img_rgb).convert("RGB")
143
+ tensor = self.transform(pil).unsqueeze(0).to(self.device)
144
+ # Wrap inference in no_grad to save memory; HAS_TORCH is
145
+ # guaranteed True here because self.model wouldn't exist
146
+ # otherwise.
147
+ with torch.no_grad():
148
+ raw = self.model(tensor).squeeze().item() # scalar in ~[1, 5]
149
+ except Exception as exc:
150
+ print(f"[BeautyAnalyzer] Inference failed: {exc}")
151
+ return self._empty_result()
152
+
153
+ # Clamp to SCUT-FBP5500's nominal 1–5 range; out-of-range outputs
154
+ # mean the regressor's extrapolating beyond its training labels.
155
+ score = max(1.0, min(5.0, float(raw)))
156
+
157
+ # Linear rescale to 0–100 for downstream display. 1.0 β†’ 0, 5.0 β†’ 100.
158
+ norm = (score - 1.0) / 4.0 * 100.0
159
+
160
+ return {
161
+ "beauty_score": round(score, 3),
162
+ "beauty_score_norm": round(norm, 1),
163
+ "beauty_model_source": self.source or "local",
164
+ }
165
+
166
+ @staticmethod
167
+ def _empty_result() -> dict[str, Any]:
168
+ return {
169
+ "beauty_score": None,
170
+ "beauty_score_norm": None,
171
+ "beauty_model_source": "unavailable",
172
+ }
analyzers/demographic_analyzer.py DELETED
@@ -1,227 +0,0 @@
1
- """
2
- DemographicAnalyzer β€” age, gender, ethnicity via three ViT classifiers.
3
-
4
- Models
5
- ------
6
- - Age : dima806/fairface_age_image_detection
7
- ViT-B/16, ~59% top-1 on FairFace 9 age buckets.
8
- - Gender : dima806/fairface_gender_image_detection
9
- ViT-B/16, ~93.4% on FairFace.
10
- - Ethnicity : cledoux42/Ethnicity_Test_v003
11
- ViT, 79.6% accuracy, macro-F1 0.797. 5-class output that
12
- we widen into the legacy 7-bucket FairFace schema so the
13
- rest of the app's distribution shape doesn't change.
14
-
15
- All three are Apache 2.0 and Hugging Face image-classification pipelines.
16
-
17
- Inputs
18
- ------
19
- img_rgb : np.ndarray (H, W, 3) uint8
20
-
21
- Outputs (dict)
22
- --------------
23
- age_range, age_estimate (softmax-weighted continuous), age_confidence,
24
- age_distribution, gender, gender_confidence, ethnicity,
25
- ethnicity_confidence, ethnicity_distribution.
26
-
27
- Notes
28
- -----
29
- The FairFace age model is a 9-bucket classifier (0-2, 3-9, …, 70+),
30
- which means the argmax bucket midpoint is always one of nine fixed
31
- numbers (24.5 for 20-29, etc.). To recover a smooth continuous estimate
32
- we compute the expected value across the full softmax β€” see
33
- ``_weighted_age_estimate``.
34
- """
35
-
36
- from typing import Any
37
-
38
- from PIL import Image
39
- from transformers import pipeline
40
-
41
-
42
- AGE_MODEL_ID = "dima806/fairface_age_image_detection"
43
- GENDER_MODEL_ID = "dima806/fairface_gender_image_detection"
44
- RACE_MODEL_ID = "cledoux42/Ethnicity_Test_v003"
45
-
46
- AGE_LABELS = ["0-2", "3-9", "10-19", "20-29", "30-39", "40-49", "50-59", "60-69", "70+"]
47
- GENDER_LABELS = ["Male", "Female"]
48
- # cledoux42 ships 5 classes (african, asian, caucasian, hispanic, indian),
49
- # but we keep the legacy 7-bucket FairFace label space internally so the
50
- # downstream distribution dict shape stays stable. Unseen buckets stay 0.
51
- RACE_LABELS = ["White", "Black", "Latino_Hispanic", "East Asian", "Southeast Asian", "Indian", "Middle Eastern"]
52
-
53
-
54
- class DemographicAnalyzer:
55
- def __init__(self):
56
- # Each classifier is a HF image-classification pipeline. They lazy
57
- # download weights from HF on first instantiation and cache them
58
- # under /root/.cache/huggingface inside the container.
59
- self.age_classifier = self._load_classifier(AGE_MODEL_ID)
60
- self.gender_classifier = self._load_classifier(GENDER_MODEL_ID)
61
- self.race_classifier = self._load_classifier(RACE_MODEL_ID)
62
-
63
- @staticmethod
64
- def _load_classifier(model_id: str):
65
- """Build one HF image-classification pipeline, logging on failure.
66
-
67
- A failed load returns None so the rest of the service continues
68
- to function and `analyze()` falls back to "unknown" demographics.
69
- """
70
- try:
71
- return pipeline("image-classification", model=model_id)
72
- except Exception as exc:
73
- print(f"[DemographicAnalyzer] Failed to load '{model_id}': {exc}")
74
- return None
75
-
76
- def analyze(self, img_rgb) -> dict[str, Any]:
77
- # Convert the numpy frame to a PIL Image once and reuse it for
78
- # all three classifier calls.
79
- pil = Image.fromarray(img_rgb)
80
-
81
- # top_k=len(labels) so we get the full softmax for each model.
82
- # We need the full age distribution to compute the weighted
83
- # expected-value age estimate.
84
- age_predictions = self._safe_predict(self.age_classifier, pil, top_k=len(AGE_LABELS))
85
- gender_predictions = self._safe_predict(self.gender_classifier, pil, top_k=2)
86
- race_predictions = self._safe_predict(self.race_classifier, pil, top_k=7)
87
-
88
- # If every classifier failed we degrade gracefully with a stub.
89
- if not age_predictions and not gender_predictions and not race_predictions:
90
- return {
91
- "age_range": "unknown",
92
- "age_estimate": 0.0,
93
- "age_confidence": 0.0,
94
- "gender": "unknown",
95
- "gender_confidence": 0.0,
96
- "ethnicity": "unknown",
97
- "ethnicity_confidence": 0.0,
98
- "age_distribution": {label: 0.0 for label in AGE_LABELS},
99
- "ethnicity_distribution": {label: 0.0 for label in RACE_LABELS},
100
- }
101
-
102
- # HF pipelines return predictions pre-sorted by score descending,
103
- # so prediction[0] is always the argmax class.
104
- age_prediction = age_predictions[0] if age_predictions else {"label": "unknown", "score": 0.0}
105
- gender_prediction = gender_predictions[0] if gender_predictions else {"label": "unknown", "score": 0.0}
106
- race_prediction = race_predictions[0] if race_predictions else {"label": "unknown", "score": 0.0}
107
-
108
- # Models occasionally return label aliases ("more than 70" instead
109
- # of "70+", "African" instead of "Black"). The normalisers map
110
- # everything back to our canonical schema.
111
- age_label = self._normalize_age_label(age_prediction["label"])
112
- gender_label = self._normalize_gender_label(gender_prediction["label"])
113
- race_label = self._normalize_race_label(race_prediction["label"])
114
-
115
- return {
116
- "age_range": age_label,
117
- "age_estimate": self._weighted_age_estimate(age_predictions),
118
- "age_confidence": round(float(age_prediction["score"]), 3),
119
- "gender": gender_label.lower(),
120
- "gender_confidence": round(float(gender_prediction["score"]), 3),
121
- "ethnicity": race_label,
122
- "ethnicity_confidence": round(float(race_prediction["score"]), 3),
123
- "age_distribution": self._distribution_map(age_predictions, self._normalize_age_label, AGE_LABELS),
124
- "ethnicity_distribution": self._distribution_map(race_predictions, self._normalize_race_label, RACE_LABELS),
125
- }
126
-
127
- @staticmethod
128
- def _normalize_age_label(label: str) -> str:
129
- """Map model output to canonical AGE_LABELS entry."""
130
- normalized = label.strip().lower()
131
- if normalized == "more than 70":
132
- return "70+"
133
- return AGE_LABELS[AGE_LABELS.index(label)] if label in AGE_LABELS else label
134
-
135
- @staticmethod
136
- def _normalize_gender_label(label: str) -> str:
137
- normalized = label.strip().lower()
138
- if normalized in {"male", "female"}:
139
- return normalized.capitalize()
140
- return label
141
-
142
- @staticmethod
143
- def _normalize_race_label(label: str) -> str:
144
- """Coalesce cledoux42's 5 classes into our 7-bucket schema."""
145
- normalized = label.strip().lower().replace("-", "_")
146
- race_aliases = {
147
- # Legacy FairFace 7-class labels
148
- "white": "White",
149
- "black": "Black",
150
- "latino_hispanic": "Latino_Hispanic",
151
- "latino hispanic": "Latino_Hispanic",
152
- "east asian": "East Asian",
153
- "southeast asian": "Southeast Asian",
154
- "indian": "Indian",
155
- "middle eastern": "Middle Eastern",
156
- # cledoux42/Ethnicity_Test_v003 5-class labels
157
- "african": "Black",
158
- "asian": "East Asian",
159
- "caucasian": "White",
160
- "hispanic": "Latino_Hispanic",
161
- }
162
- return race_aliases.get(normalized, label)
163
-
164
- # Midpoint of each FairFace age bucket β€” used as the per-bucket
165
- # "value" when we marginalise over the predicted distribution.
166
- _AGE_MIDPOINTS = {
167
- "0-2": 1.0,
168
- "3-9": 6.0,
169
- "10-19": 14.5,
170
- "20-29": 24.5,
171
- "30-39": 34.5,
172
- "40-49": 44.5,
173
- "50-59": 54.5,
174
- "60-69": 64.5,
175
- "70+": 75.0,
176
- }
177
-
178
- @classmethod
179
- def _weighted_age_estimate(cls, predictions: list[dict]) -> float:
180
- """Softmax-weighted expected age across all FairFace buckets.
181
-
182
- FairFace is a 9-bucket classifier; the argmax always snaps to one
183
- of nine fixed midpoints (24.5 for 20-29, etc.). Treating its
184
- softmax as a probability distribution and taking the expected
185
- value gives a continuous number that moves with confidence
186
- (23.1 for someone very confidently 20-29, 28.4 if some mass leaks
187
- into 30-39). Still bounded by bucket midpoints β€” true per-year
188
- accuracy would need a regression model.
189
- """
190
- total_weight = 0.0
191
- weighted_sum = 0.0
192
- for pred in predictions:
193
- label = cls._normalize_age_label(pred["label"])
194
- midpoint = cls._AGE_MIDPOINTS.get(label)
195
- if midpoint is None:
196
- continue
197
- score = float(pred["score"])
198
- weighted_sum += midpoint * score
199
- total_weight += score
200
- if total_weight == 0:
201
- return 0.0
202
- return round(weighted_sum / total_weight, 1)
203
-
204
- @classmethod
205
- def _distribution_map(cls, predictions, normalizer, all_labels):
206
- """Flatten HF predictions into {canonical_label: score} dict.
207
-
208
- Unseen labels stay at 0.0 so the shape is always all_labels-sized.
209
- """
210
- distribution = {label: 0.0 for label in all_labels}
211
- for prediction in predictions:
212
- normalized_label = normalizer(prediction["label"])
213
- if normalized_label in distribution:
214
- distribution[normalized_label] = round(float(prediction["score"]), 3)
215
- return distribution
216
-
217
- @staticmethod
218
- def _safe_predict(classifier, image, top_k: int):
219
- """Wrap classifier(...) so a single model failure can't bring
220
- down the whole demographic block."""
221
- if classifier is None:
222
- return []
223
- try:
224
- return classifier(image, top_k=top_k)
225
- except Exception as exc:
226
- print(f"[DemographicAnalyzer] Prediction failed: {exc}")
227
- return []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
analyzers/ethnicity_analyzer.py ADDED
@@ -0,0 +1,115 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ EthnicityAnalyzer β€” single-model ethnicity classifier.
3
+
4
+ Model
5
+ -----
6
+ - HF repo : cledoux42/Ethnicity_Test_v003
7
+ - Arch : Vision Transformer
8
+ - Classes : african, asian, caucasian, hispanic, indian (5)
9
+ - Reported : 79.6% accuracy, macro-F1 0.797
10
+ - License : Apache 2.0
11
+ - Source : https://huggingface.co/cledoux42/Ethnicity_Test_v003
12
+
13
+ Inputs
14
+ ------
15
+ img_rgb : np.ndarray (H, W, 3) uint8
16
+
17
+ Outputs (dict)
18
+ --------------
19
+ ethnicity : canonical label (legacy 7-bucket schema)
20
+ ethnicity_confidence : argmax softmax score
21
+ ethnicity_distribution : full {label: prob} dict, padded to all 7 buckets
22
+
23
+ Notes
24
+ -----
25
+ Split out from the old DemographicAnalyzer (which also handled age and
26
+ gender via FairFace ViTs). Age and gender now live in
27
+ InsightFaceAnalyzer; this file owns ethnicity exclusively.
28
+
29
+ The model emits 5 classes but we widen to the legacy 7-bucket FairFace
30
+ schema so the rest of the app's distribution shape stays stable.
31
+ Unseen buckets stay at 0.0.
32
+ """
33
+
34
+ from typing import Any
35
+
36
+ from PIL import Image
37
+ from transformers import pipeline
38
+
39
+
40
+ MODEL_ID = "cledoux42/Ethnicity_Test_v003"
41
+
42
+ # Legacy schema preserved from the old DemographicAnalyzer.
43
+ RACE_LABELS = [
44
+ "White", "Black", "Latino_Hispanic", "East Asian",
45
+ "Southeast Asian", "Indian", "Middle Eastern",
46
+ ]
47
+
48
+
49
+ class EthnicityAnalyzer:
50
+ def __init__(self):
51
+ self.classifier = None
52
+ try:
53
+ self.classifier = pipeline("image-classification", model=MODEL_ID)
54
+ except Exception as exc:
55
+ print(f"[EthnicityAnalyzer] Failed to load {MODEL_ID}: {exc}")
56
+
57
+ def analyze(self, img_rgb) -> dict[str, Any]:
58
+ if self.classifier is None:
59
+ return self._empty_result()
60
+
61
+ try:
62
+ pil = Image.fromarray(img_rgb)
63
+ preds = self.classifier(pil, top_k=7)
64
+ except Exception as exc:
65
+ print(f"[EthnicityAnalyzer] Prediction failed: {exc}")
66
+ return self._empty_result()
67
+
68
+ if not preds:
69
+ return self._empty_result()
70
+
71
+ # Top prediction β†’ canonical label.
72
+ top = preds[0]
73
+ top_label = self._normalize(top["label"])
74
+
75
+ distribution = {label: 0.0 for label in RACE_LABELS}
76
+ for pred in preds:
77
+ label = self._normalize(pred["label"])
78
+ if label in distribution:
79
+ distribution[label] = round(float(pred["score"]), 3)
80
+
81
+ return {
82
+ "ethnicity": top_label,
83
+ "ethnicity_confidence": round(float(top["score"]), 3),
84
+ "ethnicity_distribution": distribution,
85
+ }
86
+
87
+ @staticmethod
88
+ def _normalize(label: str) -> str:
89
+ """Map model output (5-class) to canonical 7-bucket label."""
90
+ normalized = label.strip().lower().replace("-", "_")
91
+ aliases = {
92
+ # Legacy FairFace 7-class
93
+ "white": "White",
94
+ "black": "Black",
95
+ "latino_hispanic": "Latino_Hispanic",
96
+ "latino hispanic": "Latino_Hispanic",
97
+ "east asian": "East Asian",
98
+ "southeast asian": "Southeast Asian",
99
+ "indian": "Indian",
100
+ "middle eastern": "Middle Eastern",
101
+ # cledoux42 5-class β†’ widen into 7-bucket schema
102
+ "african": "Black",
103
+ "asian": "East Asian",
104
+ "caucasian": "White",
105
+ "hispanic": "Latino_Hispanic",
106
+ }
107
+ return aliases.get(normalized, label)
108
+
109
+ @staticmethod
110
+ def _empty_result() -> dict[str, Any]:
111
+ return {
112
+ "ethnicity": "unknown",
113
+ "ethnicity_confidence": 0.0,
114
+ "ethnicity_distribution": {label: 0.0 for label in RACE_LABELS},
115
+ }
analyzers/insightface_analyzer.py ADDED
@@ -0,0 +1,171 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ InsightFaceAnalyzer β€” detection + age + gender + recognition embedding.
3
+
4
+ Model
5
+ -----
6
+ - Package : `insightface` (https://github.com/deepinsight/insightface)
7
+ - Bundle : buffalo_l (ResNet50@WebFace600K backbone, ONNX)
8
+ - Components : SCRFD-10GF detector, ArcFace 512-d recognition,
9
+ 2d106 + 3d68 landmark regressors, age + gender heads
10
+ - Size : ~280 MB (ONNX, mixed FP16/FP32)
11
+ - License : weights research-only; code Apache 2.0
12
+ - Source : https://github.com/deepinsight/insightface/tree/master/python-package
13
+
14
+ Inputs
15
+ ------
16
+ img_rgb : np.ndarray (H, W, 3) uint8
17
+
18
+ Outputs (dict)
19
+ --------------
20
+ face_bbox : [x1, y1, x2, y2] in pixel coordinates
21
+ face_confidence : SCRFD detection score
22
+ face_embedding : list[float] of length 512 (ArcFace, L2-normalised)
23
+ age_estimate : float years (regression head, not bucketed)
24
+ age_range : string bucket derived from age_estimate for
25
+ backwards compatibility with the legacy UI
26
+ gender : "male" | "female"
27
+ gender_confidence : 1.0 by default (InsightFace doesn't expose a
28
+ gender softmax score; the head is argmax-only)
29
+ _insight_landmarks_2d : list of (x, y) tuples β€” 106 points (internal)
30
+
31
+ Accuracy
32
+ --------
33
+ - Recognition (ArcFace via buffalo_l): 99.83% LFW, 96.21% IJB-B FAR=1e-4.
34
+ - Age / gender heads are widely used but lack a clean published metric.
35
+ In practice age MAE is ~5 years and gender ~94-96%.
36
+
37
+ Notes
38
+ -----
39
+ We run the bundle once per image and expose only the highest-confidence
40
+ face when multiple are detected β€” the rest of the pipeline assumes a
41
+ single subject.
42
+ """
43
+
44
+ import os
45
+ from typing import Any
46
+
47
+ import numpy as np
48
+
49
+ # insightface is a relatively heavy import; deferred so the module can
50
+ # at least load when the package isn't installed.
51
+ try:
52
+ from insightface.app import FaceAnalysis
53
+ HAS_INSIGHTFACE = True
54
+ except ImportError:
55
+ HAS_INSIGHTFACE = False
56
+
57
+
58
+ MODEL_NAME = "buffalo_l"
59
+
60
+ # Age buckets used by the legacy UI. We derive these from the regression
61
+ # output so existing screens keep working.
62
+ AGE_BUCKETS = [
63
+ (0, 3, "0-2"), (3, 10, "3-9"), (10, 20, "10-19"),
64
+ (20, 30, "20-29"), (30, 40, "30-39"), (40, 50, "40-49"),
65
+ (50, 60, "50-59"), (60, 70, "60-69"), (70, 200, "70+"),
66
+ ]
67
+
68
+
69
+ class InsightFaceAnalyzer:
70
+ def __init__(self):
71
+ self.app = None
72
+ if not HAS_INSIGHTFACE:
73
+ print(
74
+ "[InsightFaceAnalyzer] insightface package not installed; "
75
+ "detection, age, gender, and recognition will degrade to 'unknown'."
76
+ )
77
+ return
78
+
79
+ try:
80
+ # Buffalo_L bundle auto-resolves under ~/.insightface/models/.
81
+ # CPUExecutionProvider is the right default for HF Spaces;
82
+ # ctx_id=0 + 'CUDAExecutionProvider' would be the GPU path.
83
+ self.app = FaceAnalysis(
84
+ name=MODEL_NAME,
85
+ providers=["CPUExecutionProvider"],
86
+ )
87
+ # det_size=(640, 640) is the canonical SCRFD input. Smaller
88
+ # speeds inference but loses small faces.
89
+ self.app.prepare(ctx_id=-1, det_size=(640, 640))
90
+ except Exception as exc:
91
+ print(f"[InsightFaceAnalyzer] Failed to load {MODEL_NAME}: {exc}")
92
+ self.app = None
93
+
94
+ def analyze(self, img_rgb: np.ndarray) -> dict[str, Any]:
95
+ if self.app is None:
96
+ return self._empty_result()
97
+
98
+ try:
99
+ # InsightFace expects BGR, OpenCV-native order.
100
+ img_bgr = img_rgb[..., ::-1]
101
+ faces = self.app.get(img_bgr)
102
+ except Exception as exc:
103
+ print(f"[InsightFaceAnalyzer] Inference failed: {exc}")
104
+ return self._empty_result()
105
+
106
+ if not faces:
107
+ return self._empty_result()
108
+
109
+ # If multiple faces, take the highest-confidence one. The rest of
110
+ # the pipeline assumes a single subject.
111
+ face = max(faces, key=lambda f: float(f.det_score))
112
+
113
+ # Bounding box is float32 in [x1, y1, x2, y2] image-pixel space.
114
+ bbox = [float(v) for v in face.bbox.tolist()]
115
+
116
+ # Recognition embedding: 512-d, L2-normalised by InsightFace.
117
+ # Cast to plain list[float] for clean JSON.
118
+ embedding = (
119
+ [float(v) for v in face.normed_embedding.tolist()]
120
+ if getattr(face, "normed_embedding", None) is not None
121
+ else None
122
+ )
123
+
124
+ # Age head is a single float (years). Round for display.
125
+ age = float(getattr(face, "age", 0.0))
126
+
127
+ # Gender is exposed as 0 (female) / 1 (male) on Face objects.
128
+ # InsightFace doesn't surface a softmax probability β€” we report
129
+ # confidence 1.0 to indicate "argmax, no soft signal".
130
+ gender_idx = int(getattr(face, "gender", -1))
131
+ gender = "male" if gender_idx == 1 else "female" if gender_idx == 0 else "unknown"
132
+
133
+ return {
134
+ "face_bbox": bbox,
135
+ "face_confidence": round(float(face.det_score), 3),
136
+ "face_embedding": embedding,
137
+ "age_estimate": round(age, 1),
138
+ "age_range": self._bucket_age(age),
139
+ "age_confidence": 1.0,
140
+ "gender": gender,
141
+ "gender_confidence": 1.0,
142
+ # 106 2D landmarks (forehead, jaw, brows, eyes, nose, lips).
143
+ # Underscore-prefixed β†’ stripped from JSON, available to
144
+ # downstream analyzers that want a tighter face crop.
145
+ "_insight_landmarks_2d": (
146
+ [(float(p[0]), float(p[1])) for p in face.landmark_2d_106.tolist()]
147
+ if getattr(face, "landmark_2d_106", None) is not None
148
+ else None
149
+ ),
150
+ }
151
+
152
+ @staticmethod
153
+ def _bucket_age(age: float) -> str:
154
+ for lo, hi, label in AGE_BUCKETS:
155
+ if lo <= age < hi:
156
+ return label
157
+ return "unknown"
158
+
159
+ @staticmethod
160
+ def _empty_result() -> dict[str, Any]:
161
+ return {
162
+ "face_bbox": None,
163
+ "face_confidence": 0.0,
164
+ "face_embedding": None,
165
+ "age_estimate": 0.0,
166
+ "age_range": "unknown",
167
+ "age_confidence": 0.0,
168
+ "gender": "unknown",
169
+ "gender_confidence": 0.0,
170
+ "_insight_landmarks_2d": None,
171
+ }
analyzers/landmark_analyzer.py CHANGED
@@ -11,6 +11,9 @@ Model
11
  Inputs
12
  ------
13
  img_rgb : np.ndarray (H, W, 3) uint8, RGB order.
 
 
 
14
 
15
  Outputs (dict)
16
  --------------
 
11
  Inputs
12
  ------
13
  img_rgb : np.ndarray (H, W, 3) uint8, RGB order.
14
+ Always receives the full photo (not a face crop) β€” MediaPipe
15
+ has its own face detector and works best with the full
16
+ field of view to maintain z-coordinate consistency.
17
 
18
  Outputs (dict)
19
  --------------
analyzers/parsing_analyzer.py CHANGED
@@ -15,7 +15,12 @@ Model
15
 
16
  Inputs
17
  ------
18
- img_rgb : np.ndarray (H, W, 3) uint8
 
 
 
 
 
19
 
20
  Outputs (dict)
21
  --------------
 
15
 
16
  Inputs
17
  ------
18
+ img_rgb : np.ndarray (H, W, 3) uint8.
19
+ Since `app.py` v3.0, this is typically a face-cropped image
20
+ (produced by `_crop_to_face` using the InsightFace bbox)
21
+ rather than the full photo. SegFormer behaves much more
22
+ consistently when the face fills most of the input β€” the
23
+ masks become tighter and skin-stat estimates less noisy.
24
 
25
  Outputs (dict)
26
  --------------
app.py CHANGED
@@ -2,38 +2,57 @@
2
  HCP Face Analysis Microservice
3
  ==============================
4
 
5
- FastAPI service that runs seven specialized analyzers over a single photo
6
- and merges their outputs into one ~100-field facial-attribute dictionary.
 
 
7
 
8
  Pipeline (in execution order)
9
  -----------------------------
10
- 1. MediaPipe Face Landmarker 478 3D landmarks + 52 ARKit blendshapes.
11
- Produces all geometric face/eye/nose/lip/
12
- jaw features plus smiling and mouth-open.
13
-
14
- 2. DemographicAnalyzer Three ViT classifiers (FairFace age,
15
- FairFace gender, Ethnicity_Test_v003).
16
- Age is reported as a softmax-weighted
17
- continuous estimate, not a bucket midpoint.
18
-
19
- 3. ParsingAnalyzer SegFormer-B5 human parsing. Emits face
20
- and hair pixel masks plus hair length,
21
- hat detection, and skin texture/wrinkle/
22
- freckle/uniformity stats computed via
23
- OpenCV over the face mask.
24
-
25
- 4. EmotionAnalyzer HSEmotion EfficientNet-B0 8-class output
26
- plus derived valence, arousal, mood.
27
-
28
- 5. ColorAnalyzer Pixel-level LAB/HSV statistics. Reads
29
- masks from step 3 and lip/iris landmarks
30
- from step 1. No ML model.
31
-
32
- 6. ObstructionAnalyzer dima806 ViT-B/16. Glasses, sunglasses,
33
- mask flags with ~99% precision/recall.
34
-
35
- 7. HairTypeAnalyzer dima806 ViT-B/16. Curly/dreadlocks/kinky/
36
- straight/wavy at ~93% accuracy.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
  Endpoints
39
  ---------
@@ -42,15 +61,14 @@ GET /health liveness check
42
  POST /analyze multipart file upload
43
  POST /analyze-base64 JSON {"image": "<base64>"}
44
 
45
- Both POST endpoints run the same pipeline. All analyzers are lazily
46
- instantiated on first request to keep cold-start latency manageable
47
- on the Hugging Face Spaces free tier.
48
  """
49
 
50
  import os
51
- # hf_transfer gives much faster model downloads from the HF Hub on first
52
- # inference. HF_HUB_DOWNLOAD_TIMEOUT defaults to 10s which is too short
53
- # for the larger ViT checkpoints on a cold start.
54
  os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
55
  os.environ["HF_HUB_DOWNLOAD_TIMEOUT"] = "60"
56
 
@@ -64,34 +82,41 @@ from fastapi.middleware.cors import CORSMiddleware
64
  from PIL import Image
65
 
66
  from analyzers.landmark_analyzer import LandmarkAnalyzer
67
- from analyzers.demographic_analyzer import DemographicAnalyzer
68
  from analyzers.parsing_analyzer import ParsingAnalyzer
69
  from analyzers.emotion_analyzer import EmotionAnalyzer
70
  from analyzers.color_analyzer import ColorAnalyzer
71
  from analyzers.obstruction_analyzer import ObstructionAnalyzer
72
  from analyzers.hair_type_analyzer import HairTypeAnalyzer
 
 
 
73
 
74
  logging.basicConfig(level=logging.INFO)
75
  logger = logging.getLogger(__name__)
76
 
77
- app = FastAPI(title="HCP Face Analysis Service", version="2.0.0")
78
 
79
  app.add_middleware(
80
  CORSMiddleware,
81
- allow_origins=["*"], # Restrict to your domain in production
82
  allow_credentials=True,
83
  allow_methods=["*"],
84
  allow_headers=["*"],
85
  )
86
 
87
- # Analyzers are initialized lazily on first request to reduce cold-start time
 
 
88
  landmark_analyzer: Optional[LandmarkAnalyzer] = None
89
- demographic_analyzer: Optional[DemographicAnalyzer] = None
90
  parsing_analyzer: Optional[ParsingAnalyzer] = None
91
  emotion_analyzer: Optional[EmotionAnalyzer] = None
92
  color_analyzer: Optional[ColorAnalyzer] = None
93
  obstruction_analyzer: Optional[ObstructionAnalyzer] = None
94
  hair_type_analyzer: Optional[HairTypeAnalyzer] = None
 
 
95
 
96
 
97
  def _to_json_safe(value):
@@ -101,8 +126,6 @@ def _to_json_safe(value):
101
  or boolean mask logic). FastAPI's default JSON encoder doesn't
102
  handle those, so we normalise everything here before returning.
103
  """
104
- # Numpy first β€” these checks would otherwise be caught by isinstance
105
- # for dict/list because numpy.generic types are duck-typed.
106
  if isinstance(value, (np.ndarray,)):
107
  return value.tolist()
108
  if isinstance(value, (np.integer, np.floating)):
@@ -111,7 +134,6 @@ def _to_json_safe(value):
111
  return bool(value)
112
  if isinstance(value, np.generic):
113
  return value.item()
114
- # Recurse into nested containers.
115
  if isinstance(value, dict):
116
  return {str(k): _to_json_safe(v) for k, v in value.items()}
117
  if isinstance(value, (list, tuple, set)):
@@ -126,17 +148,22 @@ def get_analyzers():
126
  requests. First request pays the full model-load cost; subsequent
127
  requests are warm.
128
  """
129
- global landmark_analyzer, demographic_analyzer
130
  global parsing_analyzer, emotion_analyzer, color_analyzer
131
  global obstruction_analyzer, hair_type_analyzer
 
 
 
 
 
132
 
133
  if landmark_analyzer is None:
134
  logger.info("Loading MediaPipe Face Landmarker...")
135
  landmark_analyzer = LandmarkAnalyzer()
136
 
137
- if demographic_analyzer is None:
138
- logger.info("Loading FairFace demographics model...")
139
- demographic_analyzer = DemographicAnalyzer()
140
 
141
  if parsing_analyzer is None:
142
  logger.info("Loading SegFormer face parser...")
@@ -157,23 +184,156 @@ def get_analyzers():
157
  logger.info("Loading hair type classifier...")
158
  hair_type_analyzer = HairTypeAnalyzer()
159
 
 
 
 
 
 
 
 
160
  return (
 
161
  landmark_analyzer,
162
- demographic_analyzer,
163
  parsing_analyzer,
164
  emotion_analyzer,
165
  color_analyzer,
166
  obstruction_analyzer,
167
  hair_type_analyzer,
 
 
168
  )
169
 
170
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
171
  @app.get("/")
172
  async def root():
173
  """Service banner β€” confirms the server is reachable and which version."""
174
  return {
175
  "name": "HCP Face Analysis Service",
176
- "version": "2.0.0",
177
  "status": "running",
178
  "endpoints": {
179
  "health": "/health",
@@ -193,72 +353,15 @@ async def health():
193
  async def analyze_face(file: UploadFile = File(...)):
194
  """Multipart endpoint for direct uploads.
195
 
196
- Runs all seven analyzers and returns the merged attribute dict.
197
- See `analyze_face_base64` for the JSON-body variant the Express
198
- server calls.
199
  """
200
  try:
201
- # Decode the upload into an RGB numpy array. All analyzers
202
- # work in RGB; we don't actually need BGR but keeping it as a
203
- # local in case a future analyzer wants the OpenCV-native order.
204
  contents = await file.read()
205
  image = Image.open(io.BytesIO(contents)).convert("RGB")
206
  img_array = np.array(image)
207
-
208
- (
209
- landmarks,
210
- demographics,
211
- parsing,
212
- emotions,
213
- colors,
214
- obstructions,
215
- hair_types,
216
- ) = get_analyzers()
217
-
218
- results = {}
219
-
220
- # Step 1: MediaPipe Landmarks β†’ all geometric features + blendshapes.
221
- logger.info("Running landmark analysis...")
222
- landmark_results = landmarks.analyze(img_array)
223
- results.update(landmark_results)
224
-
225
- # Step 2: FairFace + Ethnicity ViT β†’ demographics.
226
- logger.info("Running demographic analysis...")
227
- demo_results = demographics.analyze(img_array)
228
- results.update(demo_results)
229
-
230
- # Step 3: SegFormer-B5 human parsing β†’ masks + hair length + skin stats.
231
- logger.info("Running face parsing...")
232
- parse_results = parsing.analyze(img_array)
233
- results.update(parse_results)
234
-
235
- # Step 4: HSEmotion β†’ 8-class emotion + valence/arousal/mood.
236
- logger.info("Running emotion analysis...")
237
- emo_results = emotions.analyze(img_array)
238
- results.update(emo_results)
239
-
240
- # Step 5: Pixel color analysis. Uses the face/hair masks from step 3
241
- # and MediaPipe lip/iris landmarks from step 1.
242
- logger.info("Running color analysis...")
243
- color_results = colors.analyze(
244
- img_array,
245
- skin_mask=parse_results.get("_skin_mask"),
246
- hair_mask=parse_results.get("_hair_mask"),
247
- landmarks=landmark_results.get("_raw_landmarks"),
248
- )
249
- results.update(color_results)
250
-
251
- # Step 6: ObstructionViT β†’ glasses / sunglasses / mask flags.
252
- logger.info("Running obstruction analysis...")
253
- results.update(obstructions.analyze(img_array))
254
-
255
- # Step 7: HairTypeViT β†’ curly/dreadlocks/kinky/straight/wavy.
256
- logger.info("Running hair-type analysis...")
257
- results.update(hair_types.analyze(img_array))
258
-
259
- # Remove internal fields (prefixed with underscore)
260
- results = {k: v for k, v in results.items() if not k.startswith("_")}
261
-
262
  return {"success": True, "data": _to_json_safe(results)}
263
 
264
  except Exception as e:
@@ -281,56 +384,14 @@ async def analyze_face_base64(body: dict):
281
  if not image_b64:
282
  raise HTTPException(status_code=400, detail="No image data provided")
283
 
284
- # Strip data URI prefix if present
285
  if "," in image_b64:
286
  image_b64 = image_b64.split(",", 1)[1]
287
 
288
  image_bytes = base64.b64decode(image_b64)
289
  image = Image.open(io.BytesIO(image_bytes)).convert("RGB")
290
  img_array = np.array(image)
291
-
292
- (
293
- landmarks,
294
- demographics,
295
- parsing,
296
- emotions,
297
- colors,
298
- obstructions,
299
- hair_types,
300
- ) = get_analyzers()
301
-
302
- results = {}
303
-
304
- # Same seven-step pipeline as /analyze. Kept inline (rather
305
- # than factored out) so the per-step `logger.info` cadence and
306
- # ordering stay obvious when reading either endpoint top-down.
307
- landmark_results = landmarks.analyze(img_array)
308
- results.update(landmark_results)
309
-
310
- demo_results = demographics.analyze(img_array)
311
- results.update(demo_results)
312
-
313
- parse_results = parsing.analyze(img_array)
314
- results.update(parse_results)
315
-
316
- emo_results = emotions.analyze(img_array)
317
- results.update(emo_results)
318
-
319
- color_results = colors.analyze(
320
- img_array,
321
- skin_mask=parse_results.get("_skin_mask"),
322
- hair_mask=parse_results.get("_hair_mask"),
323
- landmarks=landmark_results.get("_raw_landmarks"),
324
- )
325
- results.update(color_results)
326
-
327
- results.update(obstructions.analyze(img_array))
328
- results.update(hair_types.analyze(img_array))
329
-
330
- # Drop internal/scratch fields (leading underscore) before
331
- # returning. Keeps masks and raw landmark lists out of the JSON.
332
- results = {k: v for k, v in results.items() if not k.startswith("_")}
333
-
334
  return {"success": True, "data": _to_json_safe(results)}
335
 
336
  except HTTPException:
 
2
  HCP Face Analysis Microservice
3
  ==============================
4
 
5
+ FastAPI service that runs nine specialized analyzers over a single photo
6
+ and merges their outputs into one facial-attribute dictionary, including
7
+ a face-recognition embedding for cross-photo grouping and a numeric
8
+ "chopped score" aesthetic rating.
9
 
10
  Pipeline (in execution order)
11
  -----------------------------
12
+ 1. InsightFaceAnalyzer InsightFace buffalo_l (ONNX). SCRFD
13
+ detection + ArcFace 512-d embedding +
14
+ age regression + gender + 106 landmarks.
15
+ Replaces the previous three FairFace ViTs
16
+ and adds face matching as a new capability.
17
+
18
+ 2. LandmarkAnalyzer MediaPipe Face Landmarker. 478 3D
19
+ landmarks + 52 ARKit blendshapes β†’
20
+ geometric features, smiling, mouth_open.
21
+
22
+ 3. EthnicityAnalyzer cledoux42/Ethnicity_Test_v003 ViT.
23
+ 5-class ethnicity widened to a 7-bucket
24
+ schema for legacy compatibility.
25
+
26
+ 4. ParsingAnalyzer SegFormer-B5 human parsing. Now receives
27
+ a face-cropped image (smaller, cleaner).
28
+ Emits face/hair masks + hair length +
29
+ hat detection + OpenCV-derived skin stats.
30
+
31
+ 5. EmotionAnalyzer HSEmotion EfficientNet-B0. 8-class
32
+ emotion + valence/arousal/mood.
33
+
34
+ 6. ColorAnalyzer Pure OpenCV LAB/HSV statistics. Uses
35
+ SegFormer masks + MediaPipe lip/iris
36
+ landmarks. No ML model.
37
+
38
+ 7. ObstructionAnalyzer dima806 ViT-B/16. Glasses, sunglasses,
39
+ mask. ~99% precision on each.
40
+
41
+ 8. HairTypeAnalyzer dima806 ViT-B/16. Curly/dreadlocks/kinky/
42
+ straight/wavy. ~93% accuracy.
43
+
44
+ 9. BeautyAnalyzer Optional. ResNet-50 trained on
45
+ SCUT-FBP5500 (see training/beauty/).
46
+ Outputs a 1.0–5.0 beauty score plus a
47
+ 0–100 normalised version. Falls back to
48
+ None when no weights are loaded β€” the
49
+ AestheticAnalyzer then uses rule-based
50
+ scoring only.
51
+
52
+ 10. AestheticAnalyzer Pure-Python aggregator. Reads the merged
53
+ dict from analyzers 1–9 and produces the
54
+ final `chopped_score` (0–100, higher =
55
+ more chopped) and a per-factor breakdown.
56
 
57
  Endpoints
58
  ---------
 
61
  POST /analyze multipart file upload
62
  POST /analyze-base64 JSON {"image": "<base64>"}
63
 
64
+ All analyzers are lazily instantiated on first request to keep
65
+ cold-start latency manageable on the Hugging Face Spaces free tier.
 
66
  """
67
 
68
  import os
69
+ # hf_transfer makes initial model downloads from the HF Hub much faster.
70
+ # The default HF_HUB_DOWNLOAD_TIMEOUT (10 s) is too short for the larger
71
+ # ViT checkpoints on a cold start.
72
  os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
73
  os.environ["HF_HUB_DOWNLOAD_TIMEOUT"] = "60"
74
 
 
82
  from PIL import Image
83
 
84
  from analyzers.landmark_analyzer import LandmarkAnalyzer
85
+ from analyzers.ethnicity_analyzer import EthnicityAnalyzer
86
  from analyzers.parsing_analyzer import ParsingAnalyzer
87
  from analyzers.emotion_analyzer import EmotionAnalyzer
88
  from analyzers.color_analyzer import ColorAnalyzer
89
  from analyzers.obstruction_analyzer import ObstructionAnalyzer
90
  from analyzers.hair_type_analyzer import HairTypeAnalyzer
91
+ from analyzers.insightface_analyzer import InsightFaceAnalyzer
92
+ from analyzers.beauty_analyzer import BeautyAnalyzer
93
+ from analyzers.aesthetic_analyzer import AestheticAnalyzer
94
 
95
  logging.basicConfig(level=logging.INFO)
96
  logger = logging.getLogger(__name__)
97
 
98
+ app = FastAPI(title="HCP Face Analysis Service", version="3.0.0")
99
 
100
  app.add_middleware(
101
  CORSMiddleware,
102
+ allow_origins=["*"], # Restrict to your domain in production.
103
  allow_credentials=True,
104
  allow_methods=["*"],
105
  allow_headers=["*"],
106
  )
107
 
108
+ # Lazy slots, one per analyzer. The first request pays the full
109
+ # model-load cost; subsequent requests are warm.
110
+ insightface_analyzer: Optional[InsightFaceAnalyzer] = None
111
  landmark_analyzer: Optional[LandmarkAnalyzer] = None
112
+ ethnicity_analyzer: Optional[EthnicityAnalyzer] = None
113
  parsing_analyzer: Optional[ParsingAnalyzer] = None
114
  emotion_analyzer: Optional[EmotionAnalyzer] = None
115
  color_analyzer: Optional[ColorAnalyzer] = None
116
  obstruction_analyzer: Optional[ObstructionAnalyzer] = None
117
  hair_type_analyzer: Optional[HairTypeAnalyzer] = None
118
+ beauty_analyzer: Optional[BeautyAnalyzer] = None
119
+ aesthetic_analyzer: Optional[AestheticAnalyzer] = None
120
 
121
 
122
  def _to_json_safe(value):
 
126
  or boolean mask logic). FastAPI's default JSON encoder doesn't
127
  handle those, so we normalise everything here before returning.
128
  """
 
 
129
  if isinstance(value, (np.ndarray,)):
130
  return value.tolist()
131
  if isinstance(value, (np.integer, np.floating)):
 
134
  return bool(value)
135
  if isinstance(value, np.generic):
136
  return value.item()
 
137
  if isinstance(value, dict):
138
  return {str(k): _to_json_safe(v) for k, v in value.items()}
139
  if isinstance(value, (list, tuple, set)):
 
148
  requests. First request pays the full model-load cost; subsequent
149
  requests are warm.
150
  """
151
+ global insightface_analyzer, landmark_analyzer, ethnicity_analyzer
152
  global parsing_analyzer, emotion_analyzer, color_analyzer
153
  global obstruction_analyzer, hair_type_analyzer
154
+ global beauty_analyzer, aesthetic_analyzer
155
+
156
+ if insightface_analyzer is None:
157
+ logger.info("Loading InsightFace buffalo_l bundle...")
158
+ insightface_analyzer = InsightFaceAnalyzer()
159
 
160
  if landmark_analyzer is None:
161
  logger.info("Loading MediaPipe Face Landmarker...")
162
  landmark_analyzer = LandmarkAnalyzer()
163
 
164
+ if ethnicity_analyzer is None:
165
+ logger.info("Loading Ethnicity classifier...")
166
+ ethnicity_analyzer = EthnicityAnalyzer()
167
 
168
  if parsing_analyzer is None:
169
  logger.info("Loading SegFormer face parser...")
 
184
  logger.info("Loading hair type classifier...")
185
  hair_type_analyzer = HairTypeAnalyzer()
186
 
187
+ if beauty_analyzer is None:
188
+ logger.info("Loading beauty regressor (or no-op if untrained)...")
189
+ beauty_analyzer = BeautyAnalyzer()
190
+
191
+ if aesthetic_analyzer is None:
192
+ aesthetic_analyzer = AestheticAnalyzer()
193
+
194
  return (
195
+ insightface_analyzer,
196
  landmark_analyzer,
197
+ ethnicity_analyzer,
198
  parsing_analyzer,
199
  emotion_analyzer,
200
  color_analyzer,
201
  obstruction_analyzer,
202
  hair_type_analyzer,
203
+ beauty_analyzer,
204
+ aesthetic_analyzer,
205
  )
206
 
207
 
208
+ def _crop_to_face(img_rgb: np.ndarray, bbox, padding: float = 0.4) -> np.ndarray:
209
+ """Crop the image to a face-centred rectangle with extra context.
210
+
211
+ SegFormer and the ViT classifiers tend to do better with the face
212
+ occupying a large fraction of the input. We pad the InsightFace
213
+ bbox by `padding` (fraction of bbox size) so context like ears,
214
+ hair, and the top of the shoulders is preserved.
215
+
216
+ Returns the full image unchanged if bbox is None, malformed, or
217
+ the resulting crop would be degenerate.
218
+ """
219
+ if bbox is None or len(bbox) != 4:
220
+ return img_rgb
221
+ h, w = img_rgb.shape[:2]
222
+ try:
223
+ x1, y1, x2, y2 = bbox
224
+ bw = max(1.0, x2 - x1)
225
+ bh = max(1.0, y2 - y1)
226
+ pad_x = bw * padding
227
+ pad_y = bh * padding
228
+ cx1 = max(0, int(x1 - pad_x))
229
+ cy1 = max(0, int(y1 - pad_y))
230
+ cx2 = min(w, int(x2 + pad_x))
231
+ cy2 = min(h, int(y2 + pad_y))
232
+ if cx2 - cx1 < 32 or cy2 - cy1 < 32:
233
+ return img_rgb
234
+ return img_rgb[cy1:cy2, cx1:cx2]
235
+ except Exception:
236
+ return img_rgb
237
+
238
+
239
+ def _run_pipeline(img_array: np.ndarray) -> dict:
240
+ """Run all ten analyzers against `img_array` and return the merged dict.
241
+
242
+ Shared by /analyze and /analyze-base64. Kept as a function rather
243
+ than inlined twice so the per-step ordering is the single source
244
+ of truth.
245
+ """
246
+ (
247
+ insight,
248
+ landmarks,
249
+ ethnicities,
250
+ parsing,
251
+ emotions,
252
+ colors,
253
+ obstructions,
254
+ hair_types,
255
+ beauty,
256
+ aesthetics,
257
+ ) = get_analyzers()
258
+
259
+ results: dict = {}
260
+
261
+ # Step 1: InsightFace detection + age + gender + recognition embedding.
262
+ logger.info("Running InsightFace analysis...")
263
+ insight_results = insight.analyze(img_array)
264
+ results.update(insight_results)
265
+
266
+ # Compute a face crop once and pass it to every downstream analyzer
267
+ # that benefits from it (parsing, ethnicity, obstruction, hair type,
268
+ # beauty regressor). Falls back to the full image when InsightFace
269
+ # didn't find a face.
270
+ face_crop = _crop_to_face(img_array, insight_results.get("face_bbox"))
271
+
272
+ # Step 2: MediaPipe landmarks (works on the full image; it has its
273
+ # own internal detector).
274
+ logger.info("Running landmark analysis...")
275
+ landmark_results = landmarks.analyze(img_array)
276
+ results.update(landmark_results)
277
+
278
+ # Step 3: ethnicity classifier β€” likes a tighter face crop.
279
+ logger.info("Running ethnicity analysis...")
280
+ results.update(ethnicities.analyze(face_crop))
281
+
282
+ # Step 4: SegFormer parsing on the face crop (cleaner masks).
283
+ logger.info("Running face parsing...")
284
+ parse_results = parsing.analyze(face_crop)
285
+ results.update(parse_results)
286
+
287
+ # Step 5: HSEmotion on the face crop.
288
+ logger.info("Running emotion analysis...")
289
+ results.update(emotions.analyze(face_crop))
290
+
291
+ # Step 6: pixel-level colour analysis. Uses the face/hair masks
292
+ # from step 4 (already in face-crop coordinate space) and the
293
+ # MediaPipe lip/iris landmarks from step 2 (still in full-image
294
+ # space, normalised). We pass `face_crop` so mask coordinates
295
+ # line up; landmarks are in normalised coordinates so they map
296
+ # correctly to either image.
297
+ logger.info("Running color analysis...")
298
+ color_results = colors.analyze(
299
+ face_crop,
300
+ skin_mask=parse_results.get("_skin_mask"),
301
+ hair_mask=parse_results.get("_hair_mask"),
302
+ landmarks=landmark_results.get("_raw_landmarks"),
303
+ )
304
+ results.update(color_results)
305
+
306
+ # Step 7: obstruction classifier β€” also benefits from a face crop.
307
+ logger.info("Running obstruction analysis...")
308
+ results.update(obstructions.analyze(face_crop))
309
+
310
+ # Step 8: hair-type classifier.
311
+ logger.info("Running hair-type analysis...")
312
+ results.update(hair_types.analyze(face_crop))
313
+
314
+ # Step 9: learned beauty regressor (no-op if no weights present).
315
+ logger.info("Running beauty regressor...")
316
+ results.update(beauty.analyze(face_crop))
317
+
318
+ # Step 10: aesthetic aggregator. Reads the merged dict; no image
319
+ # input. Always runs last so it can see every other analyzer's
320
+ # outputs.
321
+ logger.info("Running aesthetic aggregator...")
322
+ results.update(aesthetics.analyze(results))
323
+
324
+ # Drop internal/scratch fields (leading underscore) before
325
+ # returning. Keeps masks and raw landmark lists out of the JSON.
326
+ results = {k: v for k, v in results.items() if not k.startswith("_")}
327
+
328
+ return results
329
+
330
+
331
  @app.get("/")
332
  async def root():
333
  """Service banner β€” confirms the server is reachable and which version."""
334
  return {
335
  "name": "HCP Face Analysis Service",
336
+ "version": "3.0.0",
337
  "status": "running",
338
  "endpoints": {
339
  "health": "/health",
 
353
  async def analyze_face(file: UploadFile = File(...)):
354
  """Multipart endpoint for direct uploads.
355
 
356
+ Runs the full ten-step pipeline and returns the merged attribute
357
+ dict. See `analyze_face_base64` for the JSON-body variant the
358
+ Express server calls.
359
  """
360
  try:
 
 
 
361
  contents = await file.read()
362
  image = Image.open(io.BytesIO(contents)).convert("RGB")
363
  img_array = np.array(image)
364
+ results = _run_pipeline(img_array)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
365
  return {"success": True, "data": _to_json_safe(results)}
366
 
367
  except Exception as e:
 
384
  if not image_b64:
385
  raise HTTPException(status_code=400, detail="No image data provided")
386
 
387
+ # Strip a possible "data:image/...;base64," prefix.
388
  if "," in image_b64:
389
  image_b64 = image_b64.split(",", 1)[1]
390
 
391
  image_bytes = base64.b64decode(image_b64)
392
  image = Image.open(io.BytesIO(image_bytes)).convert("RGB")
393
  img_array = np.array(image)
394
+ results = _run_pipeline(img_array)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
395
  return {"success": True, "data": _to_json_safe(results)}
396
 
397
  except HTTPException:
architecture.md CHANGED
@@ -2,65 +2,82 @@
2
 
3
  ## Pipeline
4
 
5
- A single photo is fed through seven analyzers. Their outputs are merged
6
- into one dictionary; later analyzers overwrite any colliding keys from
7
- earlier ones.
 
8
 
9
  ```
10
  Photo (RGB ndarray)
11
  β”‚
12
- β”œβ”€β–Ί [1] MediaPipe Face Landmarker
13
- β”‚ 478 landmarks + 52 blendshapes
14
- β”‚ β†’ all geometric features (face/eye/nose/eyebrow/lip/jaw shape),
15
- β”‚ smiling (mouthSmile blendshapes), eyes_open, possible_dimples,
16
- β”‚ possible_unibrow, facial_asymmetry_score, blendshapes dict
17
  β”‚
18
- β”œβ”€β–Ί [2] FairFace + Ethnicity ViT (DemographicAnalyzer)
19
- β”‚ β†’ age_range, age_estimate (softmax-weighted continuous), age_confidence,
20
- β”‚ gender + confidence, ethnicity + confidence, full distributions
21
  β”‚
22
- β”œβ”€β–Ί [3] SegFormer-B5 human parsing (ParsingAnalyzer)
23
- β”‚ β†’ per-class pixel masks (face, hair, hat, …)
24
- β”‚ β†’ hair_length, hair_present, hat_detected,
25
- β”‚ wrinkle_level, skin_texture_score, skin_uniformity, freckles_or_moles
26
- β”‚ (uses OpenCV stats over the SegFormer face mask for the skin rows)
27
  β”‚
28
- β”œβ”€β–Ί [4] HSEmotion EfficientNet-B0 (EmotionAnalyzer)
29
- β”‚ β†’ primary/secondary emotion, emotion_scores (8 classes),
30
- β”‚ valence, arousal, mood
31
  β”‚
32
- β”œβ”€β–Ί [5] ColorAnalyzer (no ML β€” OpenCV LAB/HSV)
33
- β”‚ inputs: SegFormer skin/hair masks + MediaPipe landmarks
 
 
 
 
 
 
 
 
 
 
34
  β”‚ β†’ skin_tone (Fitzpatrick + L*/a*/b* + hex), skin_undertone,
35
- β”‚ eye_color, hair_color (name + hex), hair_texture (pixel-Laplacian, coarse),
36
- β”‚ lip_color (shade + hex) ← lip mask built from MediaPipe outer-minus-inner lip
37
  β”‚
38
- β”œβ”€β–Ί [6] ObstructionViT β€” dima806/face_obstruction_image_detection
39
  β”‚ β†’ wearing_glasses, wearing_sunglasses, wearing_mask,
40
- β”‚ obstruction_top, obstruction_scores
 
 
 
 
41
  β”‚
42
- └─► [7] HairTypeViT β€” dima806/hair_type_image_detection
43
- β†’ hair_type (curly/dreadlocks/kinky/straight/wavy),
44
- hair_type_confidence, hair_type_scores
 
 
 
 
 
 
 
45
  ```
46
 
47
- All masks and other internal fields use a leading underscore in the key
48
- (e.g. `_skin_mask`). `app.py` strips those before returning JSON so the
49
- client never sees them.
50
 
51
  ## Attribute β†’ source map
52
 
53
- The EditProfileScreen renders only fields backed by one of these
54
- analyzers. Anything previously fed by the FaRL zero-shot classifier
55
- has been removed because its outputs were too noisy to trust.
56
-
57
  | Section | Field(s) | Source |
58
  |---|---|---|
59
- | Demographics | gender, age (continuous), age_range, ethnicity, distributions | FairFace + Ethnicity ViT |
60
- | Emotion | primary/secondary emotion, scores, valence, arousal, mood | HSEmotion |
61
- | Face Structure | face_shape (+ 4 ratios), jawline_type/angle, chin_type, cheekbone_prominence, cheek_fullness, forehead_width, facial_asymmetry_score | MediaPipe |
62
- | Hair | hair_length, hair_present | SegFormer |
63
- | Hair | hair_type (+ confidence) | HairTypeViT |
 
64
  | Hair | hair_color, hair hex | ColorAnalyzer |
65
  | Eyes | eye_shape, eye_depth, eye_spacing, eye_size, eyes_open | MediaPipe |
66
  | Eyes | eye_color | ColorAnalyzer |
@@ -70,30 +87,58 @@ has been removed because its outputs were too noisy to trust.
70
  | Lips & Mouth | lip_color (shade + hex) | ColorAnalyzer (mask from MediaPipe) |
71
  | Skin | skin_tone (Fitzpatrick, L*/a*/b*, hex), skin_undertone | ColorAnalyzer |
72
  | Skin | wrinkle_level, skin_texture_score, skin_uniformity, freckles_or_moles | SegFormer mask + OpenCV stats |
73
- | Accessories | wearing_glasses, wearing_sunglasses, wearing_mask | ObstructionViT |
74
  | Accessories | wearing_hat | SegFormer (hat class coverage) |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
75
 
76
  ## Deployment
77
 
78
- The service is built as a Docker image targeting Hugging Face Spaces
79
- free tier (2GB RAM, shared CPU). The MediaPipe `.task` is pulled at
80
- build time; all Hugging Face models lazy-download on first inference
81
- and cache under `/root/.cache/huggingface` inside the container.
 
82
 
83
  The Node/Express server forwards `/analyze-face` requests to
84
- `FACE_SERVICE_URL/analyze-base64`. The React Native client never talks
85
- to this service directly.
86
 
87
  ## Adding a new analyzer
88
 
89
- 1. Drop a new module under `analyzers/` exposing a class with
90
- `__init__()` and `analyze(img_rgb) -> dict`.
91
- 2. Import it in `app.py`, add a global slot and a lazy-load block in
92
- `get_analyzers()`, and append a `results.update(...)` call to both
93
- `/analyze` and `/analyze-base64`.
94
- 3. Surface the new keys in `client/src/screens/EditProfileScreen.js`
95
- and add a legend row in the "Analysis Method Details" section.
 
96
 
97
  Order matters: later analyzers overwrite earlier keys on collision.
98
- The specialized ViT classifiers run last so they win over any coarser
99
- signal.
 
2
 
3
  ## Pipeline
4
 
5
+ A single photo runs through ten analyzers. Their outputs are merged
6
+ into one dictionary; later analyzers can overwrite keys from earlier
7
+ ones (only intentional in a couple of places β€” `_run_pipeline` in
8
+ [app.py](app.py) is the single source of truth).
9
 
10
  ```
11
  Photo (RGB ndarray)
12
  β”‚
13
+ β”œβ”€β–Ί [1] InsightFaceAnalyzer (insightface buffalo_l, ONNX)
14
+ β”‚ β†’ face_bbox, face_confidence, face_embedding (512-d ArcFace),
15
+ β”‚ age_estimate, age_range, gender + confidences
 
 
16
  β”‚
17
+ β”œβ”€β–Ί Build face crop from face_bbox + padding. Downstream analyzers
18
+ β”‚ that benefit from a tighter input read the crop; MediaPipe gets
19
+ β”‚ the full image because it has its own detector.
20
  β”‚
21
+ β”œβ”€β–Ί [2] LandmarkAnalyzer (MediaPipe Face Landmarker)
22
+ β”‚ 478 landmarks + 52 blendshapes β†’ all geometric features,
23
+ β”‚ smiling, mouth_open (via blendshapes.jawOpen), eyes_open,
24
+ β”‚ facial_asymmetry_score, smile_asymmetry, possible_dimples,
25
+ β”‚ possible_unibrow.
26
  β”‚
27
+ β”œβ”€β–Ί [3] EthnicityAnalyzer (cledoux42/Ethnicity_Test_v003 ViT)
28
+ β”‚ β†’ ethnicity, ethnicity_confidence, ethnicity_distribution
29
+ β”‚ (cropped input).
30
  β”‚
31
+ β”œβ”€β–Ί [4] ParsingAnalyzer (SegFormer-B5 human parsing)
32
+ β”‚ β†’ _skin_mask, _hair_mask, hat_detected, hair_length,
33
+ β”‚ hair_present, wrinkle_level, skin_texture_score,
34
+ β”‚ skin_uniformity, freckles_or_moles
35
+ β”‚ (cropped input β€” cleaner masks).
36
+ β”‚
37
+ β”œβ”€β–Ί [5] EmotionAnalyzer (HSEmotion EfficientNet-B0)
38
+ β”‚ β†’ primary/secondary emotion, emotion_scores, valence,
39
+ β”‚ arousal, mood (cropped input).
40
+ β”‚
41
+ β”œβ”€β–Ί [6] ColorAnalyzer (no ML β€” OpenCV LAB/HSV)
42
+ β”‚ Reads SegFormer masks + MediaPipe lip/iris landmarks.
43
  β”‚ β†’ skin_tone (Fitzpatrick + L*/a*/b* + hex), skin_undertone,
44
+ β”‚ eye_color, hair_color (name + hex), hair_texture
45
+ β”‚ (coarse, fallback), lip_color (shade + hex)
46
  β”‚
47
+ β”œβ”€β–Ί [7] ObstructionAnalyzer (dima806/face_obstruction ViT-B/16)
48
  β”‚ β†’ wearing_glasses, wearing_sunglasses, wearing_mask,
49
+ β”‚ obstruction_scores (cropped input).
50
+ β”‚
51
+ β”œβ”€β–Ί [8] HairTypeAnalyzer (dima806/hair_type ViT-B/16)
52
+ β”‚ β†’ hair_type (curly/dreadlocks/kinky/straight/wavy),
53
+ β”‚ hair_type_confidence (cropped input).
54
  β”‚
55
+ β”œβ”€β–Ί [9] BeautyAnalyzer (ResNet-50 trained on SCUT-FBP5500)
56
+ β”‚ Optional. Loads local weights or HF Hub; if absent, output
57
+ β”‚ is None and AestheticAnalyzer falls back to rules.
58
+ β”‚ β†’ beauty_score (1.0–5.0), beauty_score_norm (0–100),
59
+ β”‚ beauty_model_source.
60
+ β”‚
61
+ └─► [10] AestheticAnalyzer (no model)
62
+ Reads the merged dict from steps 1–9 and produces the
63
+ final chopped_score (0–100) plus chopped_breakdown
64
+ showing each factor's signed contribution.
65
  ```
66
 
67
+ Internal/scratch keys use a leading underscore (`_skin_mask`,
68
+ `_hair_mask`, `_raw_landmarks`, `_insight_landmarks_2d`). `app.py`
69
+ strips them before returning JSON.
70
 
71
  ## Attribute β†’ source map
72
 
 
 
 
 
73
  | Section | Field(s) | Source |
74
  |---|---|---|
75
+ | Demographics | face_bbox, face_confidence, face_embedding (512-d), age_estimate, age_range, age_confidence, gender, gender_confidence | InsightFace buffalo_l |
76
+ | Demographics | ethnicity, ethnicity_confidence, ethnicity_distribution | EthnicityAnalyzer (cledoux42 ViT) |
77
+ | Emotion | primary/secondary emotion, emotion_scores, valence, arousal, mood | HSEmotion EffNet-B0 |
78
+ | Face Structure | face_shape (+ 4 ratios), jawline_type/angle, chin_type, cheekbone_prominence, cheek_fullness, forehead_width, facial_asymmetry_score | MediaPipe Face Landmarker |
79
+ | Hair | hair_length, hair_present | SegFormer-B5 |
80
+ | Hair | hair_type (+ confidence) | HairTypeViT (dima806) |
81
  | Hair | hair_color, hair hex | ColorAnalyzer |
82
  | Eyes | eye_shape, eye_depth, eye_spacing, eye_size, eyes_open | MediaPipe |
83
  | Eyes | eye_color | ColorAnalyzer |
 
87
  | Lips & Mouth | lip_color (shade + hex) | ColorAnalyzer (mask from MediaPipe) |
88
  | Skin | skin_tone (Fitzpatrick, L*/a*/b*, hex), skin_undertone | ColorAnalyzer |
89
  | Skin | wrinkle_level, skin_texture_score, skin_uniformity, freckles_or_moles | SegFormer mask + OpenCV stats |
90
+ | Accessories | wearing_glasses, wearing_sunglasses, wearing_mask | ObstructionViT (dima806) |
91
  | Accessories | wearing_hat | SegFormer (hat class coverage) |
92
+ | Aesthetics | beauty_score (1–5), beauty_score_norm (0–100) | BeautyAnalyzer (SCUT-FBP5500 ResNet-50) |
93
+ | Aesthetics | chopped_score (0–100), chopped_breakdown | AestheticAnalyzer (rule + learned blend) |
94
+
95
+ ## Face matching
96
+
97
+ InsightFace's ArcFace head emits a 512-d L2-normalised recognition
98
+ embedding. We store it alongside each contact in
99
+ `people.face_embedding` (pgvector). On a new photo save, the client
100
+ queries Supabase for any contact with cosine similarity β‰₯ 0.55 to the
101
+ new embedding and prompts the user *"this looks like {name}, add to
102
+ that profile?"* before creating a new contact.
103
+
104
+ LFW accuracy is 99.83%; IJB-B at FAR=1e-4 is 96.21%. For grouping
105
+ photos in a personal collection (similar lighting, same camera) this
106
+ is excellent. Identical twins and close family members can match β€” the
107
+ 0.55 threshold makes the prompt opt-in rather than auto-merge.
108
+
109
+ ## Training the beauty regressor
110
+
111
+ Live source in [training/beauty/](../training/beauty/). The script
112
+ fine-tunes a timm ResNet-50 on SCUT-FBP5500. After training, drop the
113
+ resulting `beauty_regressor.pt` into `face-service/models/` (or push
114
+ to HF Hub and set `BEAUTY_HF_REPO_ID`). `BeautyAnalyzer` picks it up
115
+ automatically on the next process boot.
116
+
117
+ Until weights exist, `beauty_score` returns None and the AestheticAnalyzer
118
+ gracefully falls back to a pure rule-based chopped score.
119
 
120
  ## Deployment
121
 
122
+ The service builds as a Docker image targeting Hugging Face Spaces
123
+ free tier (2 GB RAM, shared CPU). MediaPipe `.task` and the
124
+ InsightFace buffalo_l bundle are pulled at build time; all other
125
+ Hugging Face models lazy-download on first inference and cache under
126
+ `/root/.cache/huggingface`.
127
 
128
  The Node/Express server forwards `/analyze-face` requests to
129
+ `FACE_SERVICE_URL/analyze-base64`. The React Native client never
130
+ talks to this service directly.
131
 
132
  ## Adding a new analyzer
133
 
134
+ 1. Drop a new module under `analyzers/` with a class exposing
135
+ `__init__()` and `analyze(...) -> dict`.
136
+ 2. Import + add a lazy-load block in `app.py`'s `get_analyzers()`.
137
+ 3. Add a `results.update(...)` call inside `_run_pipeline` at the
138
+ right pipeline position.
139
+ 4. Surface the new keys in
140
+ [client/src/screens/EditProfileScreen.js](../client/src/screens/EditProfileScreen.js)
141
+ and add a legend row.
142
 
143
  Order matters: later analyzers overwrite earlier keys on collision.
144
+ The aesthetic aggregator runs last so it can see everything.
 
requirements.txt CHANGED
@@ -13,3 +13,5 @@ timm==1.0.3
13
  safetensors>=0.6.0
14
  transformers==4.45.2
15
  hsemotion>=0.2.2
 
 
 
13
  safetensors>=0.6.0
14
  transformers==4.45.2
15
  hsemotion>=0.2.2
16
+ insightface>=0.7.3
17
+ onnxruntime>=1.18.0