- Multi-stage Docker build (simplified to single-stage with pre-converted ONNX) - HTTP server with ONNX inference - API secret authentication - Uses pre-converted all-MiniLM-L6-v2 ONNX model from onnx-community - Image size: ~373 MB Generated by Mistral Vibe. Co-Authored-By: Mistral Vibe <vibe@mistral.ai>
2.6 KiB
2.6 KiB
Vector Service
A minimal Docker container that provides text embedding vectors using the all-MiniLM-L6-v2 model. The service accepts POST requests with text and returns a 384-dimensional embedding vector.
Features
- Small image size: ~373 MB (much smaller than typical Python-based solutions)
- Fast inference: Uses ONNX Runtime for efficient model execution
- API authentication: Optional API secret protection
- Pre-converted ONNX: Uses ready-to-use ONNX model from HuggingFace
Usage
Build the container
podman build -t vector-service .
Or with Docker:
docker build -t vector-service .
The build process will:
- Download the pre-converted ONNX model from
onnx-community/all-MiniLM-L6-v2-ONNXon HuggingFace - Install only the runtime dependencies (ONNX Runtime + NumPy)
- Create a minimal image (~373 MB)
Run the service
Without authentication:
podman run --rm -p 8080:8080 -d vector-service
With API secret authentication:
podman run --rm -p 8080:8080 -e API_SECRET=your-secret-key -d vector-service
Query the service
Send a POST request with JSON body:
curl -X POST http://localhost:8080/vector \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-secret-key" \
-d '{"text": "Hello World"}'
Response:
{
"vector": [0.1868536774709355, 0.8120285351760685, ...]
}
The vector has 384 dimensions.
Without authentication
If you didn't set API_SECRET, you can query without the Authorization header:
curl -X POST http://localhost:8080/vector \
-H "Content-Type: application/json" \
-d '{"text": "Hello World"}'
Files
Dockerfile: Single-stage build with pre-downloaded ONNX modelserver.py: HTTP server with ONNX inference
Technical Details
Model
- Model: all-MiniLM-L6-v2 (80 MB on disk)
- Source: Pre-converted ONNX from onnx-community/all-MiniLM-L6-v2-ONNX
- Dimensions: 384
- Format: ONNX (pre-converted)
Dependencies
- Runtime: Python 3.11, ONNX Runtime, NumPy
- No build-time dependencies needed (uses pre-converted model)
Image Size Breakdown
- Model files (ONNX + ONNX data + tokenizer + vocab): ~95 MB
- Python runtime and dependencies: ~278 MB
- Total: ~373 MB
Notes
- The build will download the pre-converted ONNX model from HuggingFace (~95 MB total)
- Much faster builds since no PyTorch or model conversion is needed
- The ONNX model includes an external data file (
model.onnx_data) which is normal for larger models - For production use, consider adding rate limiting and HTTPS