Senior Data Scientist/AI Data Engineer
Washington, District of Columbia
Job ID: 5450
CALIBRE Systems, Inc., an employee-owned mission focused solutions and digital transformation company, is looking for a highly qualified Hybrid Senior Data Scientist / AI Data Engineer to support development of AI-enabled applications and reusable data products within customer’s existing environment. This position will be on-site, hybrid, or remote, at the discretion of the customer.
The role combines hands-on retrieval-based AI engineering with data engineering, analytics, and governed data-product delivery. The individual will extend existing AI Assistant and LLM capabilities into applications, support document ingestion and retrieval, and build curated datasets, pipelines, and data products using the agency’s approved tools and access model.
This role requires expertise in retrieval-based AI, RAG, LLM integration, embeddings, vector stores, document text extraction and OCR, retrieval guardrails, retrieval ranking, source citation, AI-feature quality evaluation, Python, SQL, relational databases, data modeling, data quality, and role-based access.
The selected candidate will work closely with government staff, customer program offices, and delivery personnel to ensure delivered applications and data products meet customer specifications, acceptance criteria, role-scoped access requirements, quality thresholds, UAT expectations, security/privacy requirements, and documentation needs.
Responsibilities include, but are not limited to:
· AI Enablement and Retrieval-Based Application Support
· Extend the AI Assistant and LLM to approved data domains and document sets.
· Support case/data ingestion, OCR, embeddings, indexing, and vector-store use.
· Add Q&A, search, AI-assisted views, and source-cited answers to applications.
· Use approved AI services, retrieval guardrails, and query guardrails.
· Validate AI features against acceptance criteria, test sets, and quality thresholds.
· Data Engineering and Reusable Data Products
· Deliver curated datasets, scheduled extracts, documented pipelines, and data products.
· Support SQL, relational databases, data modeling, data quality, and role-based access.
· Build applications, dashboards, analytics, AI views, and products with Superset/Python.
· Use governed data and customer-provided catalog resources.
· Security, Access, Quality, and Delivery
· Apply role-scoped retrieval so users access only authorized data and documents.
· Support row-level security and required field-level masking.
· Implement or support access-control audit logging.
· Support UAT, defect correction, demonstrations, and production adoption.
· Prepare templates, onboarding, user/admin docs, knowledge transfer, and transition artifacts.
· Support security and ATO inputs, including SSP, assessment, and POA&M items.
Required Skills
· Python, SQL, relational databases, analytics, and application/data-product delivery.
· Data pipelines, extracts, modeling, quality, curated datasets, and reusable products.
· RAG, LLM integration, embeddings, vector stores, ranking, citations, guardrails, and validation.
· Document ingestion, text extraction, OCR, indexing, and retrieval-ready content preparation.
· RBAC, role-scoped retrieval, row security, masking, audit logging, and secure PII handling.
· Apache Superset, Python analytics, and existing Azure or authorized customer environments.
· UAT, defect correction, acceptance validation, demos, documentation, templates, and knowledge-transfer materials.
Desired Skills: (optional)
· Model Context Protocol (MCP) or similar AI tool-integration standards
· Evaluation/tracing tools such as MLflow
· PostgreSQL read replicas, CDC, dbt semantic modeling, and governed analytics/retrieval.
· CI/CD, observability, and Azure-hosted platform environments.
· Dashboards, analytics, AI-assisted views, reusable data products, or program enhancements.
· Onboarding playbooks, reusable templates, user/admin documentation, and KT materials.
· Section 508 accessibility expectations for applications and AI-enabled interfaces.
· Privacy, security, records, NIST, FedRAMP, Azure Government, and Federal ATO requirements.
· Relevant certifications: Fabric Data Engineer, dbt, Azure AI, Security+, CCSP, CDMP, or CAP.
required Experience
· Bachelor’s degree in Computer Science or a related field, or equivalent work experience.
· 2+ years building applications using retrieval-based AI and LLM services.
· 5+ years data engineering with SQL, pipelines, modeling, quality, access, and products.
· Experience integrating existing AI, RAG, and LLM capabilities into applications.
· Experience with embeddings and vector databases or vector stores.
· Experience with document text extraction and OCR.
· Experience validating AI features against criteria, test sets, or quality thresholds.
· Experience supporting UAT, defects, demos, documentation, and acceptance-ready delivery.
CALIBRE is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, disability, veteran status, age, marital status, pregnancy, genetic information, or other legally protected status. For our EEO Policy statement, please click here. If you would like more information on your EEO rights under the law, please click here.
If you would like to contact us regarding the accessibility of our website of need assistance completing the application process, please contact accommodations@calibresys.com or call 703.797.8732. Please note this contact information is for accommodations only.