CAL-003

AI Knowledge Assistant

Infrastructure as Code.
Architecture as Code.

DomainGenerative AI
ComplexityLevel 3
StatusCompleted

Overview

CAL-003 builds a governed Retrieval-Augmented Generation system over the Cloud Architect Lab engineering corpus. It retrieves project documentation and standards, generates grounded answers, and returns citations so the evidence behind a response remains inspectable.

The central engineering problem was not simply finding relevant text. It was ensuring that the system could distinguish semantically relevant content from knowledge that was appropriate to govern an answer.

Business Scenario

As the Cloud Architect Lab portfolio grows, its requirements, standards, architecture decisions, implementation notes, and validation evidence form an increasingly valuable knowledge base. CAL-003 turns that corpus into a queryable engineering assistant while preserving traceability and authority boundaries.

The assistant is intentionally read-only. It explains documented knowledge but has no authority to modify infrastructure or perform engineering actions.

Architecture

Source documents and metadata sidecars are published to Amazon S3. Amazon Bedrock Knowledge Bases manages ingestion and retrieval, Amazon Titan Text Embeddings V2 creates 1,024-dimension vector representations, and Amazon S3 Vectors provides cosine-similarity search and metadata filtering. Claude Haiku 4.5 uses the retrieved evidence to produce grounded answers with source attribution.

  • Amazon S3 for the governed source corpus
  • Amazon Bedrock Knowledge Bases for managed ingestion and retrieval
  • Amazon Titan Text Embeddings V2 for document and query embeddings
  • Amazon S3 Vectors for vector storage and similarity search
  • Claude Haiku 4.5 for grounded response generation
  • Terraform for reproducible infrastructure

Knowledge Authority

Each source carries metadata describing its case study, document type, authority class, lifecycle status, and domain. This keeps authority policy separate from the semantic embedding and allows retrieval to be constrained for a question's intent.

Requirements and standards can be treated as normative knowledge, architecture decisions as decisional knowledge, and implementation or reflective material as separate supporting classes. Semantic similarity then ranks content only within the eligible authority domain.

The Adversarial Finding

The evaluation corpus deliberately included an incorrect draft claiming that validation documentation was not required for a CAL case study. It directly contradicted the approved validation standard and was written to closely match the test question.

Incorrect Draft0.9313
Approved Standard0.8702
FindingSimilarity ≠ Authority

Vector search behaved correctly: the incorrect draft was more semantically similar and ranked first. Claude Haiku 4.5 recognized the conflict during unfiltered generation, but the architecture does not rely on the model to enforce governance. A normative-authority metadata filter excluded the draft before generation.

Semantic similarity is not knowledge authority. When a control can be enforced deterministically during retrieval, it should not depend entirely on model judgment.

Grounding, Citations, and Evaluation

Evaluation covered the complete RAG path rather than treating successful deployment as proof of system quality.

  • Knowledge publication, ingestion, and vector indexing
  • Semantic retrieval quality and ranking
  • Grounded response generation
  • Source-attribution and citation correctness
  • Metadata-filtered authority boundaries
  • Conflicting-source and adversarial behavior
  • Unsupported questions with insufficient evidence

When asked about a nonexistent CAL-008 system, the assistant returned no references and explained that the knowledge base did not contain enough evidence to answer. Refusing an unsupported answer is part of the trust model, not a failure of it.

Security Design

Security controls protect both the infrastructure and the integrity of the knowledge supplied to the model.

  • S3 Block Public Access, encryption at rest, and versioning
  • Least-privilege IAM for the Bedrock Knowledge Base role
  • Restricted access to the source bucket and S3 Vectors index
  • Embedding-model invocation limited to Titan Text Embeddings V2
  • Bedrock service trust with source-account restriction
  • Controlled corpus scope and metadata-based authority filtering
  • CloudTrail evidence for Bedrock API activity
  • No credentials or secrets stored in the knowledge corpus

Cost and Operational Fit

CAL-003 was designed for a small corpus and intermittent personal-lab usage. S3 Vectors avoided a continuously running vector database while retaining managed Bedrock integration and metadata filtering.

Observed AWS charges during the primary implementation and validation window were approximately $0.024. That figure included limited exploratory model activity outside the final Claude Haiku 4.5 path and remained subject to normal AWS billing-estimation delay.

Key Engineering Decisions

  • Use managed RAG services to focus evaluation on system behavior.
  • Select S3 Vectors for the workload's scale, cost, and filtering needs.
  • Keep authority metadata independent from semantic embeddings.
  • Apply deterministic authority constraints before generation.
  • Require citations so generated answers remain inspectable.
  • Validate deployed security controls, not only documented intent.
  • Establish reliable retrieval before granting future AI capabilities.

Lessons Learned

  • The most relevant document is not necessarily the governing source.
  • Metadata is an architectural control, not merely descriptive context.
  • Model judgment can reinforce governance but should not replace it.
  • Citations and unsupported-answer handling are essential to trust.
  • Evaluation must test retrieval, authority, generation, and security together.
  • A managed RAG system can be meaningfully validated at personal-lab cost.

Technology Stack

  • Amazon Bedrock Knowledge Bases
  • Amazon Titan Text Embeddings V2
  • Amazon S3 Vectors
  • Amazon S3
  • Claude Haiku 4.5
  • AWS Identity and Access Management
  • AWS CloudTrail
  • Terraform
  • Python and shell automation
  • Git and GitHub

Repository

The complete Terraform implementation, governed corpus workflow, architecture decisions, evaluation evidence, and security documentation are available on GitHub.

View CAL-003 on GitHub →