求人に戻る
ScovaiScovaiJobs
F

FyerX

RAG Architect

Bangalore, INオンサイト契約社員フルタイム

2026年9月23日に掲載

この求人はENで掲載されています

This is a remote position.

RAG Architect

Job Details
  • Employment Type: Contract
  • Work Mode: Remote
  • Location: Offshore
  • Total Experience Required: 6 to 10 years
  • Relevant Experience Required: 3+ years of dedicated experience designing production-grade Retrieval-Augmented Generation (RAG) architectures and optimizing LLM token throughput
  • Mandatory Certification: Google Cloud Certified Professional Cloud Database Engineer, AWS Certified Data Analytics - Specialty, or Databricks Certified Data Engineer Professional

Job Summary
We are seeking an experienced Context Window Optimization / RAG Architect to take full ownership of our enterprise generative AI retrieval performance, accuracy, and operational cost metrics. The ideal candidate will design high-throughput knowledge retrieval systems, optimize semantic context parsing, build custom re-ranking pipelines, and engineer caching grids to deliver data-grounded AI responses with minimal latency and maximum token efficiency.

Key Responsibilities
  • Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, building structures for document parsing, semantic metadata enrichment, and multi-vector lookups.
  • Optimize context window utilization patterns, designing smart parent-child chunking models, sentence-window retrievals, and sliding window strategies to eliminate irrelevant text tokens.
  • Build high-performance re-ranking layers, deploying machine learning cross-encoders (e.g., Cohere Rerank, BGE-Reranker) to score retrieved documents before feeding them into the LLM context pool.
  • Implement automated semantic caching architectures, utilizing caching layers (e.g., GPTCache) to capture recurring semantic queries, reducing API token expenditures and response latencies.
  • Establish automated data chunking pipelines, configuring ingestion routines to cleanly parse semi-structured and unstructured formats (PDFs, corporate wikis, SQL outputs) into clean vector targets.
  • Govern vector similarity spaces, fine-tuning hybrid search algorithms that cleanly combine dense semantic embeddings with sparse keyword token indexes (BM25).
  • Audit context-level hallucination rates and accuracy logs, tracking precision metrics, retrieval recall bounds, and processing speeds to systematically eliminate incorrect model generations.



Requirements

  • 6 to 10 years of enterprise data engineering, database design, or search engine engineering experience, with 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
  • Strong technical mastery of Python, vector databases (Pinecone, Milvus, Weaviate), text embedding models, open-source orchestration tools (LlamaIndex, LangChain), and SQL.
  • Deep structural understanding of context window limitations ("lost in the middle" phenomenon), multi-modal token dynamics, network data transfer speeds, and cloud memory spaces.
  • Mandatory certification: Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud vendor (AWS/GCP/Azure).

Preferred Qualifications
  • Prior experience implementing Graph RAG frameworks utilizing native knowledge graphs (e.g., Neo4j) to map complex corporate data relationship networks.
  • Familiarity with fine-tuning open-source text embedding models specifically optimized for industry-specific terminology or legacy product schemas.



職種スナップショット

職種

フルタイム

必要なスキル

Retrieval-Augmented Generation (RAG) architecture designContext window optimization and chunking strategies (parent-child, sliding windows)Designing high-throughput knowledge retrieval systems and scaling context retrieval loopsBuilding re-ranking pipelines using ML cross-encoders (e.g., Cohere Rerank, BGE-Reranker)Implementing semantic caching architectures (e.g., GPTCache) to reduce token use and latencyAutomated data ingestion and chunking for semi-structured/unstructured formats (PDFs, wikis, SQL outputs)Vector databases (Pinecone, Milvus, Weaviate)Text embedding models and fine-tuning embedding models for domain terminologyOpen-source orchestration tools for RAG (LlamaIndex, LangChain)Python programming for production-grade data and ML pipelinesSQL and database design for enterprise data engineeringHybrid search algorithm design combining dense embeddings and sparse token indexes (BM25)Monitoring and auditing retrieval accuracy, hallucination rates, precision/recall, and processing speedCloud database/analytics engineering and cloud vendor certification knowledge (AWS/GCP/Azure)Graph RAG frameworks and knowledge graph implementation (e.g., Neo4j)

似た求人

F

Microsoft Dynamics 365 F&O Solution Architect

FyerX

Bangalore, INオンサイト契約社員フルタイム
2 時間前
F

LLM Integration / LangChain Engineer

FyerX

Bangalore, INオンサイト契約社員フルタイム
2 時間前
F

ServiceNow Certified Technical Architect (CTA)

FyerX

Bangalore, INオンサイト契約社員フルタイム
2 時間前
F

Virtual Chief Information Security Officer (vCISO) / Security Director

FyerX

Bangalore, INオンサイト契約社員フルタイム
2 時間前
F

Data Engineer (Snowflake / Databricks / BigQuery)

FyerX

Bangalore, INオンサイト契約社員フルタイム
2 時間前
F

Cloud Infrastructure Architect (AWS / Azure / GCP)

FyerX

Bangalore, INオンサイト契約社員フルタイム
2 時間前
BangaloreのFyerXにおけるRAG Architect|Scovai | Scovai