F
FyerX
Vector Database Specialist (Pinecone / Milvus / Weaviate)
Bangalore South, INVor OrtVertragVollzeit
Veröffentlicht am 23. Sept. 2026
Diese Stelle wird in EN ausgeschrieben
This is a remote position.
Vector Database Specialist (Pinecone / Milvus / Weaviate)Job Details
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 4 to 8 years
- Relevant Experience Required: 2+ years of dedicated data engineering experience specializing in vector database administration, architectural index design, and high-dimensional semantic search scaling
- Mandatory Certification: Developer or Administrator certification from a major Vector DB provider (e.g., Pinecone Certified Developer, Milvus Professional) or a major cloud provider Data Engineering Specialty
Job Summary
We are seeking an experienced Vector Database Specialist to design, configure, and optimize the storage infrastructure powering our production-grade GenAI and semantic search applications. The ideal candidate will architect highly scalable vector indexes, write high-throughput embedding ingestion pipelines, configure real-time hybrid search query spaces, and maintain low-latency vector infrastructure handling millions of high-dimensional embeddings.
Key Responsibilities
- Design, deploy, and govern production vector databases (e.g., Pinecone, Milvus, Weaviate, Qdrant, or pgvector) to manage complex long-term memory structures for LLM applications.
- Build high-performance embedding ingestion pipelines, managing data chunking strategies, overlap controls, metadata schema extractions, and real-time upsert queues.
- Optimize high-dimensional vector search spaces, fine-tuning approximate nearest neighbor (ANN) graph parameters, HNSW cluster metrics, IVF index lists, and scalar quantization bounds.
- Configure advanced hybrid search architectures, engineering unified retrieval execution flows combining semantic vector lookups with traditional full-text keyword querying (BM25).
- Implement strict metadata filtering schemas, constructing optimized filter patterns to speed up context retrieval times and enforce dynamic domain isolation safety parameters.
- Monitor cluster metrics and resource optimization loops, tracking vector pod memory allocations, index reconstruction latencies, query-per-second (QPS) thresholds, and compute costs.
- Collaborate with AI and Data Engineering squads to evaluate text embedding models (e.g., OpenAI, Cohere, Hugging Face) and map vector sizing requirements cleanly to downstream application runtimes.
Requirements
- 4 to 8 years of core enterprise data engineering, database administration, or backend software development experience, with 2+ dedicated years actively scaling high-dimensional vector database frameworks.
- Strong technical mastery of Python, advanced SQL, vector similarity distance metrics (Cosine, Euclidean, Dot Product), and data transformation engines (e.g., Spark, dbt).
- Deep structural understanding of index types (HNSW, IVF, Flat), metadata index caching, memory footprint constraints, and cloud tenant auto-scaling mechanics.
- Mandatory certification: Official Vector DB specialized credential or a Professional Cloud Data Engineer certificate (AWS/GCP/Azure).
Preferred Qualifications
- Prior experience implementing real-time change data capture (CDC) architectures to automatically sync operational databases with vector catalogs.
- Familiarity with orchestration tools like LangChain, LangGraph, or LlamaIndex to structure retrieval steps for Retrieval-Augmented Generation (RAG) pipelines.
Rollenübersicht
Jobart
Vollzeit
Erforderliche Kompetenzen
Pinecone (production vector DB administration)Milvus (production vector DB administration)Weaviate (production vector DB administration)Vector index architecture (HNSW, IVF, Flat)Approximate Nearest Neighbor (ANN) tuning & quantizationEmbedding ingestion pipeline development (chunking, upserts, metadata extraction)Hybrid search engineering (vector retrieval + BM25/full-text integration)Metadata filtering and schema design for fast retrieval and domain isolationCluster monitoring and resource optimization (memory, QPS, index rebuilds, autoscaling)Embedding model evaluation and vector sizing (OpenAI, Cohere, Hugging Face)Python programmingAdvanced SQLData transformation engines (Spark, dbt)Change Data Capture (CDC) architectures for syncing operational DBs to vector catalogsRAG orchestration tools (LangChain, LangGraph, LlamaIndex)
Ähnliche Stellen
F
Enterprise AI Governance & Trust Layer Engineer
FyerX
Bangalore South, INVor OrtVertragVollzeit
vor 2 Stunden
F
MuleSoft Technical Architect
FyerX
Bangalore South, INVor OrtVertragVollzeit
vor 2 Stunden
F
Microsoft Dynamics 365 Customer Insights (Data Cloud) Architect
FyerX
Bangalore South, INVor OrtVertragVollzeit
vor 2 Stunden
F
NLP Data Scientist
FyerX
Bangalore South, INVor OrtVertragVollzeit
vor 2 Stunden
F
Agentic Workflow Engineer
FyerX
Bangalore South, INVor OrtVertragVollzeit
vor 2 Stunden
F
ServiceNow Developer (ITSM / ITOM)
FyerX
Bangalore South, INVor OrtVertragVollzeit
vor 2 Stunden