Quay lại việc làm
ScovaiScovaiJobs
F

FyerX

RAG Architect

Bangalore, INTại chỗHợp đồngToàn thời gian

Đăng 23 thg 9, 2026

Công việc này được đăng bằng EN

This is a remote position.

RAG Architect

Job Details
  • Employment Type: Contract
  • Work Mode: Remote
  • Location: Offshore
  • Total Experience Required: 6 to 10 years
  • Relevant Experience Required: 3+ years of dedicated experience designing production-grade Retrieval-Augmented Generation (RAG) architectures and optimizing LLM token throughput
  • Mandatory Certification: Google Cloud Certified Professional Cloud Database Engineer, AWS Certified Data Analytics - Specialty, or Databricks Certified Data Engineer Professional

Job Summary
We are seeking an experienced Context Window Optimization / RAG Architect to take full ownership of our enterprise generative AI retrieval performance, accuracy, and operational cost metrics. The ideal candidate will design high-throughput knowledge retrieval systems, optimize semantic context parsing, build custom re-ranking pipelines, and engineer caching grids to deliver data-grounded AI responses with minimal latency and maximum token efficiency.

Key Responsibilities
  • Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, building structures for document parsing, semantic metadata enrichment, and multi-vector lookups.
  • Optimize context window utilization patterns, designing smart parent-child chunking models, sentence-window retrievals, and sliding window strategies to eliminate irrelevant text tokens.
  • Build high-performance re-ranking layers, deploying machine learning cross-encoders (e.g., Cohere Rerank, BGE-Reranker) to score retrieved documents before feeding them into the LLM context pool.
  • Implement automated semantic caching architectures, utilizing caching layers (e.g., GPTCache) to capture recurring semantic queries, reducing API token expenditures and response latencies.
  • Establish automated data chunking pipelines, configuring ingestion routines to cleanly parse semi-structured and unstructured formats (PDFs, corporate wikis, SQL outputs) into clean vector targets.
  • Govern vector similarity spaces, fine-tuning hybrid search algorithms that cleanly combine dense semantic embeddings with sparse keyword token indexes (BM25).
  • Audit context-level hallucination rates and accuracy logs, tracking precision metrics, retrieval recall bounds, and processing speeds to systematically eliminate incorrect model generations.



Requirements

  • 6 to 10 years of enterprise data engineering, database design, or search engine engineering experience, with 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
  • Strong technical mastery of Python, vector databases (Pinecone, Milvus, Weaviate), text embedding models, open-source orchestration tools (LlamaIndex, LangChain), and SQL.
  • Deep structural understanding of context window limitations ("lost in the middle" phenomenon), multi-modal token dynamics, network data transfer speeds, and cloud memory spaces.
  • Mandatory certification: Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud vendor (AWS/GCP/Azure).

Preferred Qualifications
  • Prior experience implementing Graph RAG frameworks utilizing native knowledge graphs (e.g., Neo4j) to map complex corporate data relationship networks.
  • Familiarity with fine-tuning open-source text embedding models specifically optimized for industry-specific terminology or legacy product schemas.



Tóm tắt vai trò

Loại công việc

Toàn thời gian

Kỹ năng yêu cầu

Retrieval-Augmented Generation (RAG) architecture designContext window optimization and chunking strategies (parent-child, sliding windows)Designing high-throughput knowledge retrieval systems and scaling context retrieval loopsBuilding re-ranking pipelines using ML cross-encoders (e.g., Cohere Rerank, BGE-Reranker)Implementing semantic caching architectures (e.g., GPTCache) to reduce token use and latencyAutomated data ingestion and chunking for semi-structured/unstructured formats (PDFs, wikis, SQL outputs)Vector databases (Pinecone, Milvus, Weaviate)Text embedding models and fine-tuning embedding models for domain terminologyOpen-source orchestration tools for RAG (LlamaIndex, LangChain)Python programming for production-grade data and ML pipelinesSQL and database design for enterprise data engineeringHybrid search algorithm design combining dense embeddings and sparse token indexes (BM25)Monitoring and auditing retrieval accuracy, hallucination rates, precision/recall, and processing speedCloud database/analytics engineering and cloud vendor certification knowledge (AWS/GCP/Azure)Graph RAG frameworks and knowledge graph implementation (e.g., Neo4j)

Công việc tương tự

Evnek Technologies Pvt Ltd

Full Stack Python Developer

Evnek Technologies Pvt Ltd

Bangalore, INTại chỗHợp đồngToàn thời gian
Hôm qua
Minutes to Seconds Pty Ltd

Revit Templates & Families Creator (Remote)

Minutes to Seconds Pty Ltd

Bangalore, INTại chỗHợp đồngToàn thời gian
6 ngày trước
Evnek Technologies Pvt Ltd

Data Quality Engineer

Evnek Technologies Pvt Ltd

Bangalore, INTại chỗHợp đồngToàn thời gian
7 ngày trước
T

Technical Analyst – Tanium Endpoint Configuration Engineer

Theomnihire

Bangalore, INTại chỗHợp đồngToàn thời gian
8 ngày trước
F

Microsoft Dynamics 365 F&O Solution Architect

FyerX

Bangalore, INTại chỗHợp đồngToàn thời gian
15 ngày trước
F

LLM Integration / LangChain Engineer

FyerX

Bangalore, INTại chỗHợp đồngToàn thời gian
15 ngày trước
RAG Architect tại FyerX ở Bangalore | Scovai | Scovai