- Multi-stage Docker build (simplified to single-stage with pre-converted ONNX) - HTTP server with ONNX inference - API secret authentication - Uses pre-converted all-MiniLM-L6-v2 ONNX model from onnx-community - Image size: ~373 MB Generated by Mistral Vibe. Co-Authored-By: Mistral Vibe <vibe@mistral.ai>
100 lines
2.6 KiB
Markdown
100 lines
2.6 KiB
Markdown
# Vector Service
|
|
|
|
A minimal Docker container that provides text embedding vectors using the all-MiniLM-L6-v2 model. The service accepts POST requests with text and returns a 384-dimensional embedding vector.
|
|
|
|
## Features
|
|
|
|
- **Small image size**: ~373 MB (much smaller than typical Python-based solutions)
|
|
- **Fast inference**: Uses ONNX Runtime for efficient model execution
|
|
- **API authentication**: Optional API secret protection
|
|
- **Pre-converted ONNX**: Uses ready-to-use ONNX model from HuggingFace
|
|
|
|
## Usage
|
|
|
|
### Build the container
|
|
|
|
```bash
|
|
podman build -t vector-service .
|
|
```
|
|
|
|
Or with Docker:
|
|
|
|
```bash
|
|
docker build -t vector-service .
|
|
```
|
|
|
|
The build process will:
|
|
1. Download the pre-converted ONNX model from `onnx-community/all-MiniLM-L6-v2-ONNX` on HuggingFace
|
|
2. Install only the runtime dependencies (ONNX Runtime + NumPy)
|
|
3. Create a minimal image (~373 MB)
|
|
|
|
### Run the service
|
|
|
|
Without authentication:
|
|
```bash
|
|
podman run --rm -p 8080:8080 -d vector-service
|
|
```
|
|
|
|
With API secret authentication:
|
|
```bash
|
|
podman run --rm -p 8080:8080 -e API_SECRET=your-secret-key -d vector-service
|
|
```
|
|
|
|
### Query the service
|
|
|
|
Send a POST request with JSON body:
|
|
|
|
```bash
|
|
curl -X POST http://localhost:8080/vector \
|
|
-H "Content-Type: application/json" \
|
|
-H "Authorization: Bearer your-secret-key" \
|
|
-d '{"text": "Hello World"}'
|
|
```
|
|
|
|
Response:
|
|
```json
|
|
{
|
|
"vector": [0.1868536774709355, 0.8120285351760685, ...]
|
|
}
|
|
```
|
|
|
|
The vector has 384 dimensions.
|
|
|
|
### Without authentication
|
|
|
|
If you didn't set `API_SECRET`, you can query without the Authorization header:
|
|
|
|
```bash
|
|
curl -X POST http://localhost:8080/vector \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"text": "Hello World"}'
|
|
```
|
|
|
|
## Files
|
|
|
|
- `Dockerfile`: Single-stage build with pre-downloaded ONNX model
|
|
- `server.py`: HTTP server with ONNX inference
|
|
|
|
## Technical Details
|
|
|
|
### Model
|
|
- **Model**: all-MiniLM-L6-v2 (80 MB on disk)
|
|
- **Source**: Pre-converted ONNX from [onnx-community/all-MiniLM-L6-v2-ONNX](https://huggingface.co/onnx-community/all-MiniLM-L6-v2-ONNX)
|
|
- **Dimensions**: 384
|
|
- **Format**: ONNX (pre-converted)
|
|
|
|
### Dependencies
|
|
- Runtime: Python 3.11, ONNX Runtime, NumPy
|
|
- No build-time dependencies needed (uses pre-converted model)
|
|
|
|
### Image Size Breakdown
|
|
- Model files (ONNX + ONNX data + tokenizer + vocab): ~95 MB
|
|
- Python runtime and dependencies: ~278 MB
|
|
- Total: ~373 MB
|
|
|
|
## Notes
|
|
|
|
- The build will download the pre-converted ONNX model from HuggingFace (~95 MB total)
|
|
- Much faster builds since no PyTorch or model conversion is needed
|
|
- The ONNX model includes an external data file (`model.onnx_data`) which is normal for larger models
|
|
- For production use, consider adding rate limiting and HTTPS
|