Cross-Canon Semantic Search: System Architecture of the World English Bible RAG Database
A new Retrieval-Augmented Generation database enables thematic search across the Protestant and Catholic canons of the World English Bible. The system supports hierarchical filtering from canon level down to individual books to return relevant scripture passages.
Corpus Specification and Partitioning Schema
The underlying dataset for this Retrieval-Augmented Generation (RAG) system is composed of the World English Bible (WEB) translation. The corpus is partitioned into two major canonical configurations: the Protestant canon and the Catholic canon. By default, the system initializes queries against the Protestant canon.
- Genesis, Exodus, Leviticus, Numbers, Deuteronomy, Joshua, Judges, Ruth
- 1 Samuel, 2 Samuel, 1 Kings, 2 Kings, 1 Chronicles, 2 Chronicles, Ezra, Nehemiah, Esther
- Job, Psalms, Proverbs, Ecclesiastes, Song of Songs
- Isaiah, Jeremiah, Lamentations, Ezekiel, Daniel
- Hosea, Joel, Amos, Obadiah, Jonah, Micah, Nahum, Habakkuk, Zephaniah, Haggai, Zechariah, Malachi
- Matthew, Mark, Luke, John, Acts
- Romans, 1 Corinthians, 2 Corinthians, Galatians, Ephesians, Philippians, Colossians, 1 Thessalonians, 2 Thessalonians
- 1 Timothy, 2 Timothy, Titus, Philemon, Hebrews, James, 1 Peter, 2 Peter, 1 John, 2 John, 3 John, Jude, Revelation
Query Scoping and Hierarchical Filtering
The retrieval engine processes natural language inputs representing themes or specific topics. To optimize retrieval accuracy and reduce search space latency, the system implements a multi-tiered filtering hierarchy.
Users can scope their thematic queries by defining the active canon (Protestant or Catholic) and selecting specific books. If the book-level parameters are omitted, the query execution engine defaults to a global search across every book within the selected canon. This design allows for targeted single-book searches or cross-canon semantic sweeps.
Asynchronous Execution and UI States
The retrieval pipeline features decoupled loading states to manage network latencies during data fetching. When a thematic query is dispatched, the user interface transitions through distinct asynchronous states.
The application first displays a general text-loading state before transitioning to a dedicated scripture text-loading indicator. These granular loading states point to a multi-stage retrieval process, where the metadata index matches are processed prior to resolving and rendering the complete scripture passages in the results pane.