model select
same prompt is sent to all selected models in parallel · hover ? for model infoprompt & parameters
unsupported params are auto-mapped/stripped per model and shown as notes on each card. e.g. kimi-k3 uses max_completion_tokens & fixed sampling · deepseek accepts effort high/max only · glm-5.2 is the only glm with effort.
embedding theory
click to expandwhat are embeddings & cosine similarity?
Embeddings convert text / images / video into high-dimensional float vectors (e.g. 1024-dim). Semantically similar inputs land close together in vector space. Used for RAG retrieval, semantic search, clustering, classification, and cross-modal search (text↔image).
Cosine similarity measures the directional similarity of two vectors (magnitude is ignored):
cos(A, B) = (A · B) / (||A|| × ||B||)
- close to 1.0 — nearly identical meaning
- 0.7 ~ 0.9 — highly related (good retrieval candidate)
- 0.3 ~ 0.6 — weakly related
- close to 0 — unrelated / orthogonal
TokenHub VL embeddings are L2-normalized (||v|| = 1) by default, so a plain dot product equals cosine similarity. Check the norm values in the results below — they should be ~1.0.
Token billing: text is billed by tokens; images are split into 28×28 patches merged 2×2 (512×512 → 81 tokens, 1024×1024 → 324 tokens); video = min(seconds × fps, 64 frames) × per-frame tokens. Per-modality token counts are returned in usage.prompt_tokens_details.
benchmarks (CMTEB mean): text-4b 72.63 > text-0.6b 66.64 · MMEB-V2: vl-8b 75.26 > vl-2b 69.82
experiment setup
with 2+ inputs a cosine similarity matrix is computed. single text ≤ 2000 chars, max 128 per request (32 in this demo).
all modalities in one request are fused into a single vector (response data always has 1 entry). image 64×64 ~ 1280×1440 px, video sampled at max 64 frames.
gpt-image2 (custom-model-og-v2)
custom size: single edge ≤ 3840px, multiples of 16, aspect ratio ≤ 3:1, total pixels 655,360 ~ 8,294,400. generation can take tens of seconds.
billing notes (gpt-image2)
Pay-as-you-go (token), settled hourly. Billed at a unified rate of 10 CNY / 1M tokens; vendor-native tokens (USD) are converted with the factors below.
| type | USD / 1M tokens (vendor rate) | conversion factor |
|---|---|---|
| text input | $5 | 19.53125 |
| cached text input | $1.25 | 4.8828 |
| image input | $8 | 31.25 |
| cached image input | $2 | 7.8125 |
| image output | $30 | 117.1875 |
usage.total_tokens in the response is already in unified tokens → cost(CNY) = total_tokens ÷ 1,000,000 × 10