Compare commits

..
Author SHA1 Message Date
Kayce BasquesandGitHub 1dd8119cfd Don't process includes within code blocks (#58) 2025-12-16 14:30:38 -08:00
Jared Dillard dacabd62b4 update modules to 0.3.0 2025-12-11 12:49:49 -08:00
Jared Dillard 2f8e64cc6d Make more generic 2025-12-11 11:58:35 -08:00
Jared Dillard 71dc28e7d8 update to 0.2.0 2025-12-11 01:57:02 -08:00
Jared DillardandGitHub cbbf4b572b [docs] Use FetchContent to grab CMake modules (#61) 2025-12-11 00:55:00 -08:00
Jared DillardandGitHub 8a1a73512d It's parallel!
Updated footnote reference for sphinx-llm to clarify usage.
2025-12-10 21:58:49 -08:00
Jared Dillard 097f2c1084 add note about sphinx-llm 2025-12-08 14:13:24 -08:00
Jared DillardandGitHub 61f9b38d4d [docs] Demo all format options (#60) 2025-12-08 00:40:59 -08:00
Jared DillardandGitHub 8254002622 [docs] Improve docs wording (#59) 2025-12-07 23:26:27 -08:00
Jared Dillard 5b1b72bafb Change markdown suffix 2025-12-07 22:45:31 -08:00
Jared Dillard 3a4ccca5bb Update highlights 2025-12-06 19:38:54 -08:00
Jared DillardandGitHub 0a3da6b1f9 [docs] Add docs on the CMake workflow (#56) 2025-12-05 18:37:24 -08:00
Jared DillardandGitHub f1a3963a0d [docs] Discuss output formats in the docs (#55) 2025-12-03 13:51:09 -08:00
Jared Dillard c23c6d4468 Add tabs 2025-12-02 22:09:14 -08:00
Jared Dillard db32dc60a7 Improve CMake bits 2025-12-02 16:48:29 -08:00
Jared DillardandGitHub 543efabebb [docs] Add restbuilder builder for single page .rst builds (#53) 2025-12-01 13:34:45 -08:00
Jared Dillard 141e0e29f6 clean up the cmake 2025-10-28 12:26:45 -07:00
Jared DillardandGitHub 1b08b2f362 Add docs on markdown usage (#50) 2025-10-27 20:33:28 -07:00
Jared DillardandGitHub 5bfbb0168f Support customizable URI templates in llms.txt (#48) 2025-10-27 17:41:04 -07:00
Jared DillardandGitHub e2a80faf04 Add CMake support for also building Markdown docs in parallel (#49) 2025-10-26 20:11:09 -07:00
Jared DillardandGitHub b63801bcff Improve _sources directory handling (#47) 2025-10-13 23:53:46 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c45ebb0369 Bump the all-github-actions group with 2 updates (#45)
Bumps the all-github-actions group with 2 updates: [actions/checkout](https://github.com/actions/checkout) and [actions/setup-python](https://github.com/actions/setup-python).


Updates `actions/checkout` from 4 to 5
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v4...v5)

Updates `actions/setup-python` from 5 to 6
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v5...v6)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '5'
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: all-github-actions
- dependency-name: actions/setup-python
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: all-github-actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-10-04 18:59:29 -07:00
Jared DillardandGitHub f5dcd15889 Fix optional sphinx dependency (#44) 2025-09-17 10:51:07 -07:00
Jared DillardandGitHub e64e20133a Remove support for singlehtml (#40) 2025-08-29 15:22:32 -07:00
Jared Dillard 52949a952a Update changelog 2025-08-20 16:01:44 -07:00
Jared DillardandGitHub 3d7edbf7d9 Only allow builders that have a _sources directory (#38) 2025-08-20 15:57:47 -07:00
23 changed files with 1059 additions and 119 deletions
+5 -5
View File
@@ -10,9 +10,9 @@ jobs:
pre-commit: pre-commit:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v5
- name: Set up Python 3.10 - name: Set up Python 3.10
uses: actions/setup-python@v5 uses: actions/setup-python@v6
with: with:
python-version: "3.10" python-version: "3.10"
- uses: pre-commit/action@v3.0.1 - uses: pre-commit/action@v3.0.1
@@ -23,17 +23,17 @@ jobs:
python-version: ['3.9', '3.10', '3.11', '3.12'] python-version: ['3.9', '3.10', '3.11', '3.12']
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v5
- name: Set up Python ${{ matrix.python-version }} - name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v5 uses: actions/setup-python@v6
with: with:
python-version: ${{ matrix.python-version }} python-version: ${{ matrix.python-version }}
- name: Install dependencies - name: Install dependencies
run: | run: |
python -m pip install --upgrade pip python -m pip install --upgrade pip
pip install -e ".[dev]" pip install -e . --group dev
# - name: Run mypy # - name: Run mypy
# run: | # run: |
-7
View File
@@ -18,13 +18,6 @@ repos:
hooks: hooks:
- id: flake8 - id: flake8
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.11.2
hooks:
- id: mypy
files: ^sphinx_llms_txt/
additional_dependencies: [types-docutils]
- repo: https://github.com/sphinx-contrib/sphinx-lint - repo: https://github.com/sphinx-contrib/sphinx-lint
rev: v1.0.0 rev: v1.0.0
hooks: hooks:
+14 -11
View File
@@ -1,15 +1,18 @@
version: 2 version: 2
build: build:
os: "ubuntu-20.04" os: ubuntu-24.04
tools: tools:
python: "3.10" python: "3.13"
commands:
sphinx: - pip install cmake
configuration: docs/source/conf.py - pip install -r docs/requirements.txt
- pip install -e .
python: - cmake --workflow --preset documentation-workflow
install: # Generate llms.txt variants for demo purposes
- requirements: docs/requirements.txt - python docs/generate_llms_variants.py build/html
- method: pip # Copy built documentation to Read the Docs output directory
path: . - mkdir -p $READTHEDOCS_OUTPUT/html
- cp -r build/html/* $READTHEDOCS_OUTPUT/html/
- cp -r build/markdown/* $READTHEDOCS_OUTPUT/html/
- cp -r build/rst/* $READTHEDOCS_OUTPUT/html/
+35
View File
@@ -1,6 +1,41 @@
Changelog Changelog
========= =========
0.7.1
-----
- Don't process includes within code blocks
0.7.0
-----
- Add :confval:`llms_txt_uri_template` configuration option to control the link behavior in :confval:`llms_txt_filename`.
`#48 <https://github.com/jdillard/sphinx-llms-txt/pull/48>`_
0.6.0
-----
- Improve _sources directory handling
`#47 <https://github.com/jdillard/sphinx-llms-txt/pull/47>`_
0.5.3
-----
- Make sphinx a required dependency since there are imports from Sphinx
`#44 <https://github.com/jdillard/sphinx-llms-txt/pull/44>`_
0.5.2
-----
- Remove support for singlehtml
`#40 <https://github.com/jdillard/sphinx-llms-txt/pull/40>`_
0.5.1
-----
- Only allow builders that have _sources directory
`#38 <https://github.com/jdillard/sphinx-llms-txt/pull/38>`_
0.5.0 0.5.0
----- -----
+15
View File
@@ -0,0 +1,15 @@
cmake_minimum_required(VERSION 3.15)
project(SphinxDocs VERSION 1.0.0 LANGUAGES NONE)
# Fetch Sphinx CMake modules
include(FetchContent)
FetchContent_Declare(
sphinx_cmake_modules
GIT_REPOSITORY https://github.com/jdillard/sphinx-cmake-modules.git
GIT_TAG main
)
FetchContent_MakeAvailable(sphinx_cmake_modules)
list(APPEND CMAKE_MODULE_PATH "${sphinx_cmake_modules_SOURCE_DIR}/cmake/modules")
# Add documentation
add_subdirectory(docs)
+53
View File
@@ -0,0 +1,53 @@
{
"version": 6,
"configurePresets": [
{
"name": "documentation",
"displayName": "Documentation Build",
"description": "Configure project with documentation environment setup",
"binaryDir": "${sourceDir}/build"
}
],
"buildPresets": [
{
"name": "html",
"displayName": "Build HTML Documentation",
"configurePreset": "documentation",
"targets": ["html"]
},
{
"name": "markdown",
"displayName": "Build Markdown Documentation",
"configurePreset": "documentation",
"targets": ["markdown"]
},
{
"name": "rst",
"displayName": "Build reStructuredText Documentation",
"configurePreset": "documentation",
"targets": ["rst"]
},
{
"name": "docs-parallel",
"displayName": "Build all output formats in parallel",
"configurePreset": "documentation",
"targets": ["html", "markdown", "rst"]
}
],
"workflowPresets": [
{
"name": "documentation-workflow",
"displayName": "Documentation Build Workflow",
"steps": [
{
"type": "configure",
"name": "documentation"
},
{
"type": "build",
"name": "docs-parallel"
}
]
}
]
}
+7
View File
@@ -0,0 +1,7 @@
include(SphinxUtils)
setup_sphinx_environment()
add_sphinx_builder(html)
add_sphinx_builder(markdown)
add_sphinx_builder(rst)
+59
View File
@@ -0,0 +1,59 @@
#!/usr/bin/env python3
"""
Generate variant llms.txt files for demo purposes.
Takes the generated llms.txt (with _sources links) and creates:
- llms.txt - default with _sources links (unchanged)
- llms.md.txt - .html.md links
- llms.rst.txt - .rst links
"""
import re
import sys
from pathlib import Path
def get_base_url() -> str:
"""Import base_url from conf.py."""
sys.path.insert(0, str(Path(__file__).parent / "source"))
from conf import html_baseurl # noqa: E402
return html_baseurl
def generate_variants(build_dir: Path) -> None:
"""Generate llms.txt variants from the original file."""
original = build_dir / "llms.txt"
if not original.exists():
print(f"Error: {original} not found")
sys.exit(1)
content = original.read_text()
base_url = get_base_url()
# Pattern to match links like: https://.../_sources/{docname}.rst.txt
link_pattern = re.compile(
rf"({re.escape(base_url)})_sources/([a-zA-Z0-9_/\-]+)\.rst\.txt"
)
# Generate .html.md variant
md_content = link_pattern.sub(r"\1\2.html.md", content)
(build_dir / "llms.md.txt").write_text(md_content)
print(f"Generated: {build_dir / 'llms.md.txt'} (.html.md links)")
# Generate .rst variant
rst_content = link_pattern.sub(r"\1\2.rst", content)
(build_dir / "llms.rst.txt").write_text(rst_content)
print(f"Generated: {build_dir / 'llms.rst.txt'} (.rst links)")
print(f"Kept: {build_dir / 'llms.txt'} (_sources links)")
if __name__ == "__main__":
if len(sys.argv) > 1:
build_dir = Path(sys.argv[1])
else:
build_dir = Path("build/html")
generate_variants(build_dir)
+4
View File
@@ -2,5 +2,9 @@ furo
esbonio esbonio
sphinx-contributors sphinx-contributors
sphinx sphinx
sphinx-design
sphinx-llms-txt sphinx-llms-txt
sphinx-inline-tabs
sphinxext-opengraph sphinxext-opengraph
sphinx-markdown-builder
sphinxcontrib-restbuilder
+173
View File
@@ -261,6 +261,178 @@ If you want to include absolute URLs for resources in your documentation, you ca
When this option is set, all resolved paths in directives will be prefixed with this URL, creating absolute paths in the generated files. When this option is set, all resolved paths in directives will be prefixed with this URL, creating absolute paths in the generated files.
.. _customizing_uri_links:
Customizing URI Links in llms.txt
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
By default, the ``llms.txt`` file links to source files in the ``_sources`` directory when available, falling back to HTML pages when sources aren't available.
You can customize this behavior using URI templates with :confval:`llms_txt_uri_template`:
.. code-block:: python
# Default: Link to source files, if _sources exists
llms_txt_uri_template = "{base_url}_sources/{docname}{suffix}{sourcelink_suffix}"
# Default: Link to HTML pages instead, if _sources doesn't exist
llms_txt_uri_template = "{base_url}{docname}.html"
# Manual: Link to a custom markdown build
llms_txt_uri_template = "{base_url}{docname}.md"
.. _available_template_variables:
Available Template Variables
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Your URI template can use the following variables:
- ``{base_url}`` - The base URL from ``html_baseurl`` configuration (includes trailing slash)
- ``{docname}`` - The document name (e.g., ``index``, ``guide/intro``)
- ``{suffix}`` - The source file suffix (e.g., ``.rst``, ``.md``) - may be empty if no source file exists
- ``{sourcelink_suffix}`` - The suffix from ``html_sourcelink_suffix`` configuration (e.g., ``.txt``)
.. tip::
Instead of using the default of linking to ``_sources``, you can generate Markdown and/or reStructuredText files from your documentation and link to those in ``llms.txt``.
See :ref:`cmake_workflow` for an example of building both HTML and Markdown and/or reStructuredText in parallel.
Note that ``_sources`` is still needed for ``llms-full.txt`` at this time.
.. _cmake_workflow:
CMake Workflow
^^^^^^^^^^^^^^
This project uses CMake to orchestrate documentation builds across multiple output formats, serving as a simple demo of the functionality.
This approach enables parallel builds and integrates well with CI/CD platforms like Read the Docs.
Building multiple formats allows you to compare what works best for your docs, as well as allows users to choose which format to feed to their LLM.
Use :confval:`llms_txt_uri_template` to configure links to point to your preferred format.
Key Files
~~~~~~~~~
These configuration files serve as a simple example of a Sphinx site hosted on Read The Docs, some modification may be needed.
.. code-block:: text
.
├── .readthedocs.yml
├── CMakeLists.txt
├── CMakePresets.json
└── docs/
└── CMakeLists.txt
Each section below contains a summary of the file's purpose, the full contents of the file, and a table describing key lines that may need modification.
.. dropdown:: .readthedocs.yml
:chevron: down-up
A Read The Docs config file that installs dependencies, then runs the full documentation workflow which builds all output formats in parallel, and copies them into a single deploy location.
.. literalinclude:: ../../.readthedocs.yml
:language: yaml
:lines: 1-9,11,14-
:linenos:
:emphasize-lines: 9, 14-15
.. list-table::
:header-rows: 1
:width: 100%
:widths: 15 85
* - Line
- Description
* - **9**
- Update the path if your requirements file is in a different location
* - **13-14**
- Modify the copy commands for the output formats you deploy
.. dropdown:: CMakeLists.txt
:chevron: down-up
A CMake config file that sets up the project, fetches the shared `sphinx-cmake-modules <https://github.com/jdillard/sphinx-cmake-modules>`_, and includes the docs subdirectory.
.. literalinclude:: ../../CMakeLists.txt
:language: cmake
:linenos:
:emphasize-lines: 9, 15
.. list-table::
:header-rows: 1
:width: 100%
:widths: 15 85
* - Line
- Description
* - **9**
- Update the ``GIT_TAG`` to use a different version or commit hash
* - **15**
- Change if your docs subdirectory has a different location
.. dropdown:: docs/CMakeLists.txt
:chevron: down-up
A CMake config file that includes the `SphinxUtils <https://github.com/jdillard/sphinx-cmake-modules/blob/v0.1.0/SphinxUtils.cmake>`_ module from FetchContent and defines the documentation-specific build targets.
.. literalinclude:: ../CMakeLists.txt
:language: cmake
:linenos:
:emphasize-lines: 5-7
.. list-table::
:header-rows: 1
:width: 100%
:widths: 15 85
* - Line
- Description
* - **5-7**
- Add or remove calls based on which output formats you need
.. dropdown:: CMakePresets.json
:chevron: down-up
Defines presets for configuring and building documentation:
- **Configure Presets:** Sets up the build directory.
- **Build Presets:** Defines build formats individually and all in parallel.
- **Workflow Presets:** Runs the configure preset followed by the parallel build preset.
.. literalinclude:: ../../CMakePresets.json
:language: json
:linenos:
:emphasize-lines: 18-23, 24-29, 34
.. list-table::
:header-rows: 1
:width: 100%
:widths: 15 85
* - Line
- Description
* - **18-23**
- Remove this preset to disable Markdown documentation builds
* - **24-29**
- Remove this preset to disable reStructuredText documentation builds
* - **34**
- Modify the targets list to build only the output formats you need in parallel
Usage
~~~~~
To build documentation locally using CMake:
.. code-block:: console
# Run the full workflow (configure + build all formats)
cmake --workflow --preset documentation-workflow
# Or configure and build separately
cmake --preset documentation
cmake --build --preset html # Build HTML only
cmake --build --preset docs-parallel # Build all formats
.. _integration_examples: .. _integration_examples:
Integration Examples Integration Examples
@@ -285,6 +457,7 @@ Here's a complete example showing multiple :doc:`configuration-values`:
This is a comprehensive documentation set for our project. This is a comprehensive documentation set for our project.
It includes API references, usage examples, and tutorials. It includes API references, usage examples, and tutorials.
""" """
llms_txt_uri_template = "{base_url}{docname}.md"
# Path handling # Path handling
html_baseurl = "https://docs.example.com/" html_baseurl = "https://docs.example.com/"
+10 -1
View File
@@ -15,12 +15,17 @@ import subprocess
project = "sphinx-llms-txt" project = "sphinx-llms-txt"
copyright = "Jared Dillard" copyright = "Jared Dillard"
author = "Jared Dillard" author = "Jared Dillard"
llms_txt_code_files = ["+:../../sphinx_llms_txt/*.py"] llms_txt_code_files = ["+:../../sphinx_llms_txt/*.py"]
llms_txt_summary = """ llms_txt_summary = """
A Sphinx extension that generates a summary llms.txt file,written in Markdown, A Sphinx extension that generates a summary llms.txt file,written in Markdown,
and a single combined documentation llms-full.txt file, written in reStructuredText. and a single combined documentation llms-full.txt file, written in reStructuredText.
""" """
# This doesn't seem to be supported
# rst_file_suffix = ".html.rst"
markdown_file_suffix = ".html.md"
# check if the current commit is tagged as a release (vX.Y.Z) # check if the current commit is tagged as a release (vX.Y.Z)
try: try:
GIT_TAG_OUTPUT = subprocess.check_output(["git", "tag", "--points-at", "HEAD"]) GIT_TAG_OUTPUT = subprocess.check_output(["git", "tag", "--points-at", "HEAD"])
@@ -49,6 +54,10 @@ extensions = [
"sphinx.ext.intersphinx", "sphinx.ext.intersphinx",
"sphinx_contributors", "sphinx_contributors",
"sphinx_llms_txt", "sphinx_llms_txt",
"sphinxcontrib.restbuilder",
"sphinx_inline_tabs",
"sphinx.ext.extlinks",
"sphinx_design",
] ]
# The language for content autogenerated by Sphinx. Refer to documentation # The language for content autogenerated by Sphinx. Refer to documentation
@@ -88,7 +97,7 @@ html_theme_options = {
"source_directory": "docs/source/", "source_directory": "docs/source/",
} }
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/" html_baseurl = "https://sphinx-llms-txt.readthedocs.org/en/latest/"
# -- Options for HTMLHelp output --------------------------------------------- # -- Options for HTMLHelp output ---------------------------------------------
+9
View File
@@ -58,6 +58,15 @@ Project Configuration Values
.. versionadded:: 0.2.0 .. versionadded:: 0.2.0
.. confval:: llms_txt_uri_template
- **Type**: string or ``None``
- **Default**: ``None``
- **Description**: Template string for generating URIs in ``llms.txt``.
See :ref:`customizing_uri_links`.
.. versionadded:: 0.7.0
.. confval:: llms_txt_directives .. confval:: llms_txt_directives
- **Type**: list of strings - **Type**: list of strings
+1 -1
View File
@@ -19,7 +19,7 @@ Local development
.. code-block:: console .. code-block:: console
pip install -e ".[dev]" pip install -e . --group dev
#. Install pre-commit Git hook scripts: #. Install pre-commit Git hook scripts:
+59 -5
View File
@@ -4,15 +4,17 @@ Getting Started
Installation Installation
------------ ------------
Directly install via ``pip`` by using: Directly install by using:
.. tab:: via pip
.. code-block:: bash .. code-block:: bash
pip install sphinx-llms-txt pip install sphinx-llms-txt
Or with ``conda`` via ``conda-forge``: .. tab:: via conda:
.. code:: .. code-block:: bash
conda install -c conda-forge sphinx-llms-txt conda install -c conda-forge sphinx-llms-txt
@@ -32,7 +34,59 @@ After the HTML finishes building, **sphinx-llms-txt** will output the location o
sphinx-llms-txt: Created /path/to/_build/html/llms-full.txt with 45 sources and 6879 lines sphinx-llms-txt: Created /path/to/_build/html/llms-full.txt with 45 sources and 6879 lines
sphinx-llms-txt: created /path/to/_build/html/llms.txt sphinx-llms-txt: created /path/to/_build/html/llms.txt
.. _choosing-output-format:
.. tip:: Make sure to confirm the accuracy of the output files after installs and upgrades. Choosing an Output Format
-------------------------
By default, **sphinx-llms-txt** requires no additional configuration and links to raw reStructuredText source files created by the HTML builder.
For optimal LLM support, see the alternative builders below and the :ref:`CMake workflow <cmake_workflow>` for setup.
.. list-table:: Output Format Comparison
:header-rows: 1
:widths: 18 27 27 27
* -
- Default
- Markdown
- reStructuredText
* - **Setup**
- No config
- CMake [#sphinxllm]_
- CMake
* - **Builder**
- Native [#native]_
- `sphinx-markdown-builder`_
- `sphinxcontrib-restbuilder`_
* - **Format**
- Raw reStructuredText source
- Rendered Markdown [#rendered]_
- Rendered reStructuredText [#rendered]_
* - **LLM Readability**
- Good - preserves structure for simple syntax
- Excellent - native LLM format
- Good - Can provide more structured content
* - **Key Advantage**
- Zero setup required
- More compact (less input tokens)
- Can preserve Sphinx semantics
* - **Key Disadvantage**
- Raw directives won't be parsed [#autodoc]_
- Loses structure from complex directives
- Can lose structure from complex directives
* - **llms-full.txt support**
- Supported with above caveats
- Pending `support <https://github.com/liran-funaro/sphinx-markdown-builder/pull/37>`__ [#pending]_
- Pending `support <https://github.com/sphinx-contrib/restbuilder/pull/35>`__ [#pending]_
.. _sphinx-markdown-builder: https://pypi.org/project/sphinx-markdown-builder/
.. _sphinxcontrib-restbuilder: https://pypi.org/project/sphinxcontrib-restbuilder/
.. rubric:: Footnotes
.. [#sphinxllm] See `sphinx-llm <https://github.com/jacobtomlinson/sphinx-llm>`_ as an alternative for CMake-free Markdown builds.
.. [#native] Uses raw :confval:`_sources/ <sphinx:html_copy_source>` files created by Sphinx's HTML builder with some minor enhancements.
.. [#autodoc] Directives like ``autodoc`` will appear as raw directive syntax rather than the extracted docstrings.
.. [#pending] PRs that add ``llms-full.txt`` concatenation support have yet to be released.
.. [#rendered] Directives are expanded and processed before output, so content like autodoc docstrings will be included.
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
+13 -9
View File
@@ -8,21 +8,23 @@ A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Mar
Demo Demo
---- ----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example. This Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as an example of the default output format.
Alternative :ref:`output formats <choosing-output-format>` are also available. For example: `Markdown`_ and `reStructuredText`_.
Highlights Highlights
---------- ----------
1. **Content Collection**: Quickly gathers content from _sources, without needing a separate build **Zero Configuration**
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content Add the extension to your ``conf.py`` and you're done.
3. **Path Resolution**: Transforms relative paths in directives to full paths The extension automatically collects your documentation and generates both ``llms.txt`` and ``llms-full.txt`` during your normal Sphinx build.
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown **Intelligent Content Processing**
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText Automatically resolves ``include`` directives, transforms relative paths, and handles your documentation structure without manual intervention.
5. **Content Filtering**: Allows you to exclude specific pages or sections **Customizable When Needed**
6. **Source Code**: Allows you to include specific source code files Filter content, include source code files, or integrate with alternative output formats like Markdown for even better LLM compatibility.
See :doc:`getting-started` for output format options and :doc:`configuration-values` for all settings.
.. toctree:: .. toctree::
:maxdepth: 2 :maxdepth: 2
@@ -36,6 +38,8 @@ Highlights
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt .. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt .. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
.. _Markdown: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.md.txt
.. _reStructuredText: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.rst.txt
.. _Sphinx: http://sphinx-doc.org/ .. _Sphinx: http://sphinx-doc.org/
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg .. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
+6 -2
View File
@@ -26,13 +26,16 @@ classifiers = [
license = {text = "MIT"} license = {text = "MIT"}
readme = "README.md" readme = "README.md"
dynamic = ["version"] dynamic = ["version"]
dependencies = [
"sphinx",
]
[project.urls] [project.urls]
download = "https://pypi.org/project/sphinx-llms-txt/" download = "https://pypi.org/project/sphinx-llms-txt/"
source = "https://github.com/jdillard/sphinx-llms-txt" source = "https://github.com/jdillard/sphinx-llms-txt"
changelog = "https://github.com/jdillard/sphinx-llms-txt/blob/master/CHANGELOG.rst" changelog = "https://github.com/jdillard/sphinx-llms-txt/blob/master/CHANGELOG.rst"
[project.optional-dependencies] [dependency-groups]
dev = [ dev = [
"pytest>=7.0.0", "pytest>=7.0.0",
"black", "black",
@@ -40,12 +43,13 @@ dev = [
"mypy", "mypy",
"isort", "isort",
"pre-commit", "pre-commit",
"sphinx",
] ]
test = [ test = [
"pytest>=7.0.0", "pytest>=7.0.0",
] ]
[tool.setuptools]
packages = ["sphinx_llms_txt"]
[tool.setuptools.dynamic] [tool.setuptools.dynamic]
version = {attr = "sphinx_llms_txt.__version__"} version = {attr = "sphinx_llms_txt.__version__"}
+13 -6
View File
@@ -21,7 +21,7 @@ from .manager import LLMSFullManager
from .processor import DocumentProcessor from .processor import DocumentProcessor
from .writer import FileWriter from .writer import FileWriter
__version__ = "0.5.0" __version__ = "0.7.1"
# Export classes needed by tests # Export classes needed by tests
__all__ = [ __all__ = [
@@ -85,6 +85,7 @@ def build_finished(app: Sphinx, exception):
config = { config = {
"llms_txt_file": app.config.llms_txt_file, "llms_txt_file": app.config.llms_txt_file,
"llms_txt_filename": app.config.llms_txt_filename, "llms_txt_filename": app.config.llms_txt_filename,
"llms_txt_uri_template": app.config.llms_txt_uri_template,
"llms_txt_title": app.config.llms_txt_title, "llms_txt_title": app.config.llms_txt_title,
"llms_txt_summary": summary, "llms_txt_summary": summary,
"llms_txt_full_file": app.config.llms_txt_full_file, "llms_txt_full_file": app.config.llms_txt_full_file,
@@ -107,15 +108,15 @@ def build_finished(app: Sphinx, exception):
_manager.update_page_title(docname, title) _manager.update_page_title(docname, title)
# Create the combined file # Create the combined file
_manager.combine_sources(str(app.outdir), str(app.srcdir)) _manager.combine_sources(app.outdir, app.srcdir)
def setup(app: Sphinx) -> Dict[str, Any]: def setup(app: Sphinx) -> Dict[str, Any]:
"""Set up the Sphinx extension.""" """Set up the Sphinx extension."""
# Add configuration options
app.add_config_value("llms_txt_file", True, "env") app.add_config_value("llms_txt_file", True, "env")
app.add_config_value("llms_txt_filename", "llms.txt", "env") app.add_config_value("llms_txt_filename", "llms.txt", "env")
app.add_config_value("llms_txt_uri_template", None, "env")
app.add_config_value("llms_txt_full_file", True, "env") app.add_config_value("llms_txt_full_file", True, "env")
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env") app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
app.add_config_value("llms_txt_full_max_size", None, "env") app.add_config_value("llms_txt_full_max_size", None, "env")
@@ -127,15 +128,21 @@ def setup(app: Sphinx) -> Dict[str, Any]:
app.add_config_value("llms_txt_code_files", [], "env") app.add_config_value("llms_txt_code_files", [], "env")
app.add_config_value("llms_txt_code_base_path", None, "env") app.add_config_value("llms_txt_code_base_path", None, "env")
# Connect to Sphinx events def builder_inited(app):
app.connect("doctree-resolved", doctree_resolved) """Used to limit what builders are allowed to run the extension."""
app.connect("build-finished", build_finished)
allowed_builders = ["html", "dirhtml"]
if hasattr(app, "builder") and app.builder.name in allowed_builders:
# Reset manager and root paragraph for each build # Reset manager and root paragraph for each build
global _manager, _root_first_paragraph global _manager, _root_first_paragraph
_manager = LLMSFullManager() _manager = LLMSFullManager()
_root_first_paragraph = "" _root_first_paragraph = ""
app.connect("doctree-resolved", doctree_resolved)
app.connect("build-finished", build_finished)
app.connect("builder-inited", builder_inited)
return { return {
"version": __version__, "version": __version__,
"parallel_read_safe": True, "parallel_read_safe": True,
+15 -22
View File
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
""" """
import fnmatch import fnmatch
from typing import Any, Dict, List, Optional, Tuple from typing import Any, Dict, List, Tuple
from sphinx.environment import BuildEnvironment from sphinx.environment import BuildEnvironment
from sphinx.util import logging from sphinx.util import logging
@@ -16,8 +16,8 @@ class DocumentCollector:
def __init__(self): def __init__(self):
self.page_titles: Dict[str, str] = {} self.page_titles: Dict[str, str] = {}
self.master_doc: Optional[str] = None self.master_doc: str = None
self.env: Optional[BuildEnvironment] = None self.env: BuildEnvironment = None
self.config: Dict[str, Any] = {} self.config: Dict[str, Any] = {}
self.app = None self.app = None
@@ -60,7 +60,7 @@ class DocumentCollector:
else: else:
return [source_suffix] # String format return [source_suffix] # String format
def _get_docname_suffix(self, docname: str, sources_dir) -> Optional[str]: def _get_docname_suffix(self, docname: str, sources_dir) -> str:
""" """
Determine the source suffix for a given docname by checking which Determine the source suffix for a given docname by checking which
file exists. file exists.
@@ -102,7 +102,7 @@ class DocumentCollector:
return None return None
def get_page_order(self, sources_dir=None) -> List[Tuple[str, Optional[str]]]: def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
"""Get the correct page order from the toctree structure. """Get the correct page order from the toctree structure.
Args: Args:
@@ -114,7 +114,7 @@ class DocumentCollector:
if not self.env or not self.master_doc: if not self.env or not self.master_doc:
return [] return []
page_order: List[Tuple[str, Optional[str]]] = [] page_order = []
visited = set() visited = set()
def collect_from_toctree(docname: str): def collect_from_toctree(docname: str):
@@ -126,7 +126,7 @@ class DocumentCollector:
# Add the current document with its suffix # Add the current document with its suffix
if docname not in [doc for doc, _ in page_order]: if docname not in [doc for doc, _ in page_order]:
suffix: Optional[str] = None suffix = None
if sources_dir: if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir) suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix)) page_order.append((docname, suffix))
@@ -135,21 +135,18 @@ class DocumentCollector:
try: try:
# Look for toctree_includes which contains the direct children # Look for toctree_includes which contains the direct children
if ( if (
self.env hasattr(self.env, "toctree_includes")
and hasattr(self.env, "toctree_includes")
and docname in self.env.toctree_includes and docname in self.env.toctree_includes
): ):
for child_docname in self.env.toctree_includes[docname]: for child_docname in self.env.toctree_includes[docname]:
collect_from_toctree(str(child_docname)) collect_from_toctree(child_docname)
# Try to use dependencies to find related documents # Try to use dependencies to find related documents
elif ( elif (
self.env hasattr(self.env, "dependencies")
and hasattr(self.env, "dependencies")
and docname in self.env.dependencies and docname in self.env.dependencies
): ):
# Extract the dependent documents from the dependencies dict # Extract the dependent documents from the dependencies dict
for child_docname_obj in self.env.dependencies[docname]: for child_docname in self.env.dependencies[docname]:
child_docname = str(child_docname_obj)
# Only add documents actually in the document set # Only add documents actually in the document set
if ( if (
hasattr(self.env, "all_docs") hasattr(self.env, "all_docs")
@@ -157,11 +154,7 @@ class DocumentCollector:
): ):
collect_from_toctree(child_docname) collect_from_toctree(child_docname)
# Fallback to titles or other available references # Fallback to titles or other available references
elif ( elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
self.env
and hasattr(self.env, "titles")
and hasattr(self.env, "all_docs")
):
# Get all document names # Get all document names
all_docnames = list(self.env.all_docs.keys()) all_docnames = list(self.env.all_docs.keys())
@@ -192,7 +185,7 @@ class DocumentCollector:
] ]
) )
for docname in remaining: for docname in remaining:
suffix: Optional[str] = None suffix = None
if sources_dir: if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir) suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix)) page_order.append((docname, suffix))
@@ -200,8 +193,8 @@ class DocumentCollector:
return page_order return page_order
def filter_excluded_pages( def filter_excluded_pages(
self, page_order: List[Tuple[str, Optional[str]]] self, page_order: List[Tuple[str, str]]
) -> List[Tuple[str, Optional[str]]]: ) -> List[Tuple[str, str]]:
"""Filter out excluded pages from the page order.""" """Filter out excluded pages from the page order."""
exclude_patterns = self.config.get("llms_txt_exclude") exclude_patterns = self.config.get("llms_txt_exclude")
if exclude_patterns: if exclude_patterns:
+47 -27
View File
@@ -5,7 +5,7 @@ Main manager module for sphinx-llms-txt.
import glob import glob
import subprocess import subprocess
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple, Union, cast from typing import Any, Dict, List, Optional, Tuple, Union
from sphinx.application import Sphinx from sphinx.application import Sphinx
from sphinx.environment import BuildEnvironment from sphinx.environment import BuildEnvironment
@@ -150,8 +150,8 @@ class LLMSFullManager:
self.ignored_pages.add(docname) self.ignored_pages.add(docname)
def _filter_ignored_pages( def _filter_ignored_pages(
self, page_order: Union[List[str], List[Tuple[str, Optional[str]]]] self, page_order: Union[List[str], List[Tuple[str, str]]]
) -> Union[List[str], List[Tuple[str, Optional[str]]]]: ) -> Union[List[str], List[Tuple[str, str]]]:
"""Filter out ignored pages from page_order.""" """Filter out ignored pages from page_order."""
filtered_pages = [] filtered_pages = []
for item in page_order: for item in page_order:
@@ -164,7 +164,7 @@ class LLMSFullManager:
if docname not in self.ignored_pages: if docname not in self.ignored_pages:
filtered_pages.append(item) filtered_pages.append(item)
return cast(Union[List[str], List[Tuple[str, Optional[str]]]], filtered_pages) return filtered_pages
def set_config(self, config: Dict[str, Any]): def set_config(self, config: Dict[str, Any]):
"""Set configuration options.""" """Set configuration options."""
@@ -197,7 +197,6 @@ class LLMSFullManager:
possible_sources = [ possible_sources = [
Path(outdir) / "_sources", Path(outdir) / "_sources",
Path(outdir) / "html" / "_sources", Path(outdir) / "html" / "_sources",
Path(outdir) / "singlehtml" / "_sources",
] ]
for path in possible_sources: for path in possible_sources:
@@ -205,27 +204,45 @@ class LLMSFullManager:
sources_dir = path sources_dir = path
break break
if not sources_dir: # Get the correct page order (with or without source suffixes)
logger.warning(
"Could not find _sources directory, skipping llms-full creation"
)
return
# Get the correct page order with source suffixes
page_order = self.collector.get_page_order(sources_dir) page_order = self.collector.get_page_order(sources_dir)
if not page_order: if not page_order:
logger.warning( logger.warning("Could not determine page order, skipping file generation")
"Could not determine page order, skipping llms-full creation"
)
return return
# Apply exclusion filter if configured # Apply exclusion filter if configured
page_order = self.collector.filter_excluded_pages(page_order) page_order = self.collector.filter_excluded_pages(page_order)
# Determine output file name and location # If no sources directory, only generate llms.txt and return early
if not sources_dir:
# Generate llms.txt if requested
if self.config.get("llms_txt_file"):
filtered_page_order = self._filter_ignored_pages(page_order)
self.writer.write_verbose_info_to_file(
filtered_page_order,
self.collector.page_titles,
0, # No line count since no llms-full.txt
sources_dir,
)
# Only warn if user explicitly wants llms-full.txt
if self.config.get("llms_txt_full_file"):
# Check if html_copy_source is False
if self.app and not self.app.config.html_copy_source:
logger.warning(
"Could not find _sources directory, skipping llms-full.txt."
"Set html_copy_source = True in conf.py to enable."
)
else:
logger.warning(
"Could not find _sources directory, skipping llms-full.txt"
)
return
# Determine output file name and location for llms-full.txt
output_filename = self.config.get("llms_txt_full_filename") output_filename = self.config.get("llms_txt_full_filename")
output_path = Path(outdir) / str(output_filename) output_path = Path(outdir) / output_filename
# Log discovered files and page order # Log discovered files and page order
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}") logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
@@ -286,7 +303,7 @@ class LLMSFullManager:
content_parts = [] content_parts = []
# Track code files for later processing # Track code files for later processing
code_file_parts: List[str] = [] code_file_parts = []
# Count lines in code files (initially 0) # Count lines in code files (initially 0)
code_files_line_count = 0 code_files_line_count = 0
@@ -365,7 +382,7 @@ class LLMSFullManager:
if not (size_limit_exceeded and should_abort_early): if not (size_limit_exceeded and should_abort_early):
# Get all source files in the _sources directory using configured suffixes # Get all source files in the _sources directory using configured suffixes
source_suffixes = self._get_source_suffixes() source_suffixes = self._get_source_suffixes()
all_source_files: List[Path] = [] all_source_files = []
for src_suffix in source_suffixes: for src_suffix in source_suffixes:
# Avoid duplicate extensions when source_suffix == source_link_suffix # Avoid duplicate extensions when source_suffix == source_link_suffix
if src_suffix == source_link_suffix: if src_suffix == source_link_suffix:
@@ -507,6 +524,7 @@ class LLMSFullManager:
filtered_page_order, filtered_page_order,
self.collector.page_titles, self.collector.page_titles,
total_line_count, total_line_count,
sources_dir,
) )
return return
elif action == "note": elif action == "note":
@@ -520,6 +538,7 @@ class LLMSFullManager:
filtered_page_order, filtered_page_order,
self.collector.page_titles, self.collector.page_titles,
total_line_count, total_line_count,
sources_dir,
) )
return return
elif action == "keep": elif action == "keep":
@@ -538,7 +557,10 @@ class LLMSFullManager:
if success and self.config.get("llms_txt_file"): if success and self.config.get("llms_txt_file"):
filtered_page_order = self._filter_ignored_pages(page_order) filtered_page_order = self._filter_ignored_pages(page_order)
self.writer.write_verbose_info_to_file( self.writer.write_verbose_info_to_file(
filtered_page_order, self.collector.page_titles, total_line_count filtered_page_order,
self.collector.page_titles,
total_line_count,
sources_dir,
) )
def _read_source_file(self, file_path: Path, docname: str) -> Tuple[str, int]: def _read_source_file(self, file_path: Path, docname: str) -> Tuple[str, int]:
@@ -735,9 +757,9 @@ class LLMSFullManager:
title = Path(title_str[len(base_path) :]) title = Path(title_str[len(base_path) :])
except ValueError: except ValueError:
# File is not relative to srcdir, use filename # File is not relative to srcdir, use filename
title = Path(file_path.name) title = file_path.name
else: else:
title = Path(file_path.name) title = file_path.name
# Format as code block with equals underline # Format as code block with equals underline
title_str = str(title) title_str = str(title)
@@ -769,9 +791,7 @@ class LLMSFullManager:
return code_parts, sorted(processed_files) return code_parts, sorted(processed_files)
def _create_code_files_section_header( def _create_code_files_section_header(self, file_paths: List[Path] = None) -> str:
self, file_paths: Optional[List[Path]] = None
) -> str:
"""Create the section header for source code files. """Create the section header for source code files.
Args: Args:
@@ -819,7 +839,7 @@ class LLMSFullManager:
return "" return ""
# Convert to relative paths if possible and create tree structure # Convert to relative paths if possible and create tree structure
tree_data: Dict[str, Any] = {} tree_data = {}
for file_path in sorted(file_paths): for file_path in sorted(file_paths):
# Get relative path from source directory for display # Get relative path from source directory for display
@@ -872,7 +892,7 @@ class LLMSFullManager:
current[parts[-1]] = None # None indicates it's a file current[parts[-1]] = None # None indicates it's a file
# Convert tree structure to string representation # Convert tree structure to string representation
lines: List[str] = [] lines = []
self._format_tree_node(tree_data, lines, "", True) self._format_tree_node(tree_data, lines, "", True)
# Indent each line for reStructuredText code block # Indent each line for reStructuredText code block
+82 -2
View File
@@ -128,9 +128,12 @@ class DocumentProcessor:
Returns: Returns:
Processed content with directive paths properly resolved Processed content with directive paths properly resolved
""" """
# Get code block ranges to skip directives inside them
code_block_ranges = self._get_code_block_ranges(content)
# Get the configured path directives to process # Get the configured path directives to process
default_path_directives = ["image", "figure"] default_path_directives = ["image", "figure", "literalinclude"]
custom_path_directives = self.config.get("llms_txt_directives") or [] custom_path_directives = self.config.get("llms_txt_directives")
path_directives = set(default_path_directives + custom_path_directives) path_directives = set(default_path_directives + custom_path_directives)
# Build the regex pattern to match all configured directives # Build the regex pattern to match all configured directives
@@ -143,6 +146,11 @@ class DocumentProcessor:
is_test = "pytest" in str(source_path) and "subdir" in str(source_path) is_test = "pytest" in str(source_path) and "subdir" in str(source_path)
def replace_directive_path(match, base_url=base_url, is_test=is_test): def replace_directive_path(match, base_url=base_url, is_test=is_test):
# Check if this directive is within a code block
if self._is_in_code_block(match.start(), code_block_ranges):
# This directive is inside a code block, don't process it
return match.group(0)
prefix = match.group(1) # The entire directive prefix including whitespace prefix = match.group(1) # The entire directive prefix including whitespace
path = match.group(3).strip() # The path argument path = match.group(3).strip() # The path argument
@@ -276,6 +284,71 @@ class DocumentProcessor:
return possible_paths return possible_paths
def _get_code_block_ranges(self, content: str) -> List[Tuple[int, int]]:
"""Find all code block ranges in the content.
Args:
content: The source content to analyze
Returns:
List of (start, end) tuples representing code block character
ranges
"""
code_block_ranges = []
# Match code block as well as `code` and `sourcecode` aliases
code_block_pattern = re.compile(
r"^(\s*)\.\.\s+(code-block|code|sourcecode)::\s*\S*\s*$", re.MULTILINE
)
for match in code_block_pattern.finditer(content):
start_pos = match.start()
indent = match.group(1)
indent_len = len(indent)
# Find the end of the code block by looking for the next line
# that is not indented more than the directive
block_start = match.end()
pos = block_start
# Skip any blank lines immediately after the directive
while pos < len(content) and content[pos] in "\n":
pos += 1
# Find where the code block ends
lines = content[pos:].split("\n")
block_end = pos
for line in lines:
if line.strip(): # Non-empty line
# Check indentation level
line_indent = len(line) - len(line.lstrip())
if line_indent <= indent_len:
# The block ends when we find a line that is indented
# less than the directive itself
break
block_end += len(line) + 1 # +1 for the newline
code_block_ranges.append((start_pos, block_end))
return code_block_ranges
def _is_in_code_block(
self, match_start: int, code_block_ranges: List[Tuple[int, int]]
) -> bool:
"""Check if a match position is within a code block.
Args:
match_start: The starting position of the match
code_block_ranges: List of (start, end) tuples for code blocks
Returns:
True if the match is within a code block, False otherwise
"""
for block_start, block_end in code_block_ranges:
if block_start <= match_start < block_end:
return True
return False
def _process_includes(self, content: str, source_path: Path) -> str: def _process_includes(self, content: str, source_path: Path) -> str:
"""Process include directives in content. """Process include directives in content.
@@ -286,11 +359,18 @@ class DocumentProcessor:
Returns: Returns:
Processed content with include directives replaced with included content Processed content with include directives replaced with included content
""" """
code_block_ranges = self._get_code_block_ranges(content)
# Find all include directives using regex # Find all include directives using regex
include_pattern = build_directive_pattern(["include"]) include_pattern = build_directive_pattern(["include"])
# Function to replace each include with content # Function to replace each include with content
def replace_include(match): def replace_include(match):
# Check if this include is within a code block
if self._is_in_code_block(match.start(), code_block_ranges):
# This include is inside a code block, don't process it
return match.group(0)
include_path = match.group(3) include_path = match.group(3)
directive_part = match.group( directive_part = match.group(
1 1
+68 -12
View File
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
""" """
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple, Union from typing import Any, Dict, List, Tuple, Union
from sphinx.application import Sphinx from sphinx.application import Sphinx
from sphinx.util import logging from sphinx.util import logging
@@ -14,16 +14,47 @@ logger = logging.getLogger(__name__)
class FileWriter: class FileWriter:
"""Handles writing processed content to output files.""" """Handles writing processed content to output files."""
def __init__( def __init__(self, config: Dict[str, Any], outdir: str = None, app: Sphinx = None):
self,
config: Dict[str, Any],
outdir: Optional[str] = None,
app: Optional[Sphinx] = None,
):
self.config = config self.config = config
self.outdir = outdir self.outdir = outdir
self.app = app self.app = app
def _resolve_uri_template(self, sources_dir: Path = None) -> str:
"""Resolve which URI template to use based on configuration and sources_dir.
Args:
sources_dir: Path to _sources directory (None if not found)
Returns:
The template string to use for generating URIs
"""
# If custom template exists
custom_template = self.config.get("llms_txt_uri_template")
if custom_template:
# Validate user's template by checking for valid variable names
try:
# Try formatting with test valid values to validate syntax
test_values = {
"base_url": "http://example.com/",
"docname": "test",
"suffix": ".rst",
"sourcelink_suffix": ".txt",
}
custom_template.format(**test_values)
return custom_template
except (KeyError, ValueError) as e:
logger.warning(
f"sphinx-llms-txt: Invalid llms_txt_uri_template: {e}. "
f"Falling back to default."
)
# Else, use one of the default templates
if sources_dir:
return "{base_url}_sources/{docname}{suffix}{sourcelink_suffix}"
else:
return "{base_url}{docname}.html"
def write_combined_file( def write_combined_file(
self, content_parts: List[str], output_path: Path, total_line_count: int self, content_parts: List[str], output_path: Path, total_line_count: int
) -> bool: ) -> bool:
@@ -52,9 +83,10 @@ class FileWriter:
def write_verbose_info_to_file( def write_verbose_info_to_file(
self, self,
page_order: Union[List[str], List[Tuple[str, Optional[str]]]], page_order: Union[List[str], List[Tuple[str, str]]],
page_titles: Dict[str, str], page_titles: Dict[str, str],
total_line_count: int = 0, total_line_count: int = 0,
sources_dir: Path = None,
) -> bool: ) -> bool:
"""Write summary information to the llms.txt file. """Write summary information to the llms.txt file.
@@ -62,6 +94,7 @@ class FileWriter:
page_order: Ordered list of document names or (docname, suffix) tuples page_order: Ordered list of document names or (docname, suffix) tuples
page_titles: Dictionary mapping docnames to titles page_titles: Dictionary mapping docnames to titles
total_line_count: Total number of lines in the combined content total_line_count: Total number of lines in the combined content
sources_dir: Path to _sources directory (None if not found)
Returns: Returns:
True if successful, False otherwise True if successful, False otherwise
@@ -72,13 +105,13 @@ class FileWriter:
) )
return False return False
output_path = Path(self.outdir) / str(self.config.get("llms_txt_filename")) output_path = Path(self.outdir) / self.config.get("llms_txt_filename")
try: try:
with open(output_path, "w", encoding="utf-8") as f: with open(output_path, "w", encoding="utf-8") as f:
project_name = "llms-txt Summary" project_name = "llms-txt Summary"
# First priority: use title from config if available # First priority: use title from config if available
if self.config.get("llms_txt_title"): if self.config.get("llms_txt_title"):
project_name = str(self.config.get("llms_txt_title")) project_name = self.config.get("llms_txt_title")
# Second priority: use project name from Sphinx app if available # Second priority: use project name from Sphinx app if available
elif ( elif (
self.app self.app
@@ -107,14 +140,37 @@ class FileWriter:
if not base_url.endswith("/"): if not base_url.endswith("/"):
base_url += "/" base_url += "/"
# Get sourcelink suffix from Sphinx config
sourcelink_suffix = ""
if self.app and hasattr(self.app.config, "html_sourcelink_suffix"):
sourcelink_suffix = self.app.config.html_sourcelink_suffix
# Handle empty string case specially
if sourcelink_suffix == "":
sourcelink_suffix = "" # Keep it empty
elif not sourcelink_suffix.startswith("."):
sourcelink_suffix = "." + sourcelink_suffix
# Resolve which template to use
uri_template = self._resolve_uri_template(sources_dir)
for item in page_order: for item in page_order:
# Handle both old format (str) and new format (tuple) # Handle both old format (str) and new format (tuple)
if isinstance(item, tuple): if isinstance(item, tuple):
docname, _ = item docname, suffix = item
else: else:
docname = item docname = item
suffix = None
title = page_titles.get(docname, docname) title = page_titles.get(docname, docname)
f.write(f"- [{title}]({base_url}{docname}.html)\n")
uri = uri_template.format(
base_url=base_url,
docname=docname,
suffix=suffix or "",
sourcelink_suffix=sourcelink_suffix,
)
f.write(f"- [{title}]({uri})\n")
logger.info(f"sphinx-llms-txt: created {output_path}") logger.info(f"sphinx-llms-txt: created {output_path}")
return True return True
+187
View File
@@ -41,6 +41,42 @@ def test_setup_returns_valid_dict():
assert "parallel_write_safe" in result assert "parallel_write_safe" in result
def test_builder_inited_with_disallowed_builder():
"""Test that disallowed builders do not trigger extension setup."""
import sphinx_llms_txt
# Reset global state
sphinx_llms_txt._manager = sphinx_llms_txt.LLMSFullManager()
sphinx_llms_txt._root_first_paragraph = ""
# Mock a Sphinx app with a disallowed builder
class MockBuilder:
name = "text" # Not in allowed list
class MockApp:
def __init__(self):
self.config_values = {}
self.connections = {}
self.builder = MockBuilder()
def add_config_value(self, name, default, rebuild):
self.config_values[name] = (default, rebuild)
def connect(self, event, handler):
self.connections[event] = handler
app = MockApp()
setup(app)
# Trigger builder-inited
builder_inited_handler = app.connections["builder-inited"]
builder_inited_handler(app)
# With disallowed builder, other events should NOT be connected
assert "doctree-resolved" not in app.connections
assert "build-finished" not in app.connections
def test_document_collector_initialization(): def test_document_collector_initialization():
"""Test initialization of DocumentCollector.""" """Test initialization of DocumentCollector."""
collector = DocumentCollector() collector = DocumentCollector()
@@ -162,6 +198,31 @@ def test_process_includes(tmp_path):
assert processed_content == expected_content assert processed_content == expected_content
def test_process_includes_in_code_block(tmp_path):
"""Test that an `include` within a `code-block` is not processed."""
# Create a processor
config = {"llms_txt_directives": []}
processor = DocumentProcessor(config)
# Create a source file that uses include syntax within a `code-block`
source_content = (
"Normal paragraph.\n\n"
".. code-block:: rst\n\n"
" .. include:: foo.txt\n\n"
"Another normal paragraph."
)
source_file = tmp_path / "source.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Run the include directive processor
processed_content = processor._process_includes(source_content, source_file)
# Check that the include directive was not processed
expected_content = source_content
assert processed_content == expected_content
def test_process_includes_with_relative_paths(tmp_path): def test_process_includes_with_relative_paths(tmp_path):
"""Test that include directives with relative paths are processed correctly.""" """Test that include directives with relative paths are processed correctly."""
# Create a processor # Create a processor
@@ -758,6 +819,7 @@ def test_summary_default_uses_first_paragraph():
llms_txt_summary = None # Not configured llms_txt_summary = None # Not configured
llms_txt_file = True llms_txt_file = True
llms_txt_filename = "llms.txt" llms_txt_filename = "llms.txt"
llms_txt_uri_template = None
llms_txt_title = None llms_txt_title = None
llms_txt_full_file = True llms_txt_full_file = True
llms_txt_full_filename = "llms-full.txt" llms_txt_full_filename = "llms-full.txt"
@@ -995,3 +1057,128 @@ def test_code_files_ignored_patterns(tmp_path, caplog):
assert ( assert (
"Code file pattern 'docs/**/*.rst' ignored." in captured_warnings[0] "Code file pattern 'docs/**/*.rst' ignored." in captured_warnings[0]
), f"Warning message should contain expected text. Got: {captured_warnings[0]}" ), f"Warning message should contain expected text. Got: {captured_warnings[0]}"
def test_llms_txt_generated_without_sources_dir(tmp_path):
"""Test that llms.txt is generated even when _sources directory doesn't exist."""
from sphinx_llms_txt.manager import LLMSFullManager
# Create manager
manager = LLMSFullManager()
# Set config to enable llms.txt
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_full_file": True,
"llms_txt_full_filename": "llms-full.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
manager.set_config(config)
# Create directories (but no _sources)
outdir = tmp_path / "build"
srcdir = tmp_path / "source"
outdir.mkdir()
srcdir.mkdir()
# Mock env with documents
class MockEnv:
all_docs = {"index": None, "about": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda self: "Home"})(),
"about": type("TitleNode", (), {"astext": lambda self: "About"})(),
}
toctree_includes = {"index": ["about"]}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Update page titles directly in the collector
manager.update_page_title("index", "Home")
manager.update_page_title("about", "About")
# Call combine_sources - should generate llms.txt even without _sources
manager.combine_sources(str(outdir), str(srcdir))
# Verify llms.txt was created
llms_txt = outdir / "llms.txt"
assert llms_txt.exists(), "llms.txt should be generated even without _sources"
# Verify llms-full.txt was NOT created (since no _sources)
llms_full_txt = outdir / "llms-full.txt"
assert (
not llms_full_txt.exists()
), "llms-full.txt should not be generated without _sources"
# Read llms.txt and verify it has content
with open(llms_txt, "r", encoding="utf-8") as f:
content = f.read()
# Should contain page titles and links
assert "Home" in content
assert "About" in content
assert "index.html" in content
assert "about.html" in content
def test_llms_txt_no_warning_when_full_file_disabled(tmp_path, caplog):
"""
Test that no warning is logged when llms_txt_full_file=False and
_sources doesn't exist.
"""
from unittest.mock import patch
from sphinx_llms_txt.manager import LLMSFullManager
# Create manager
manager = LLMSFullManager()
# Set config with llms_txt_full_file=False
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_full_file": False, # User doesn't want llms-full.txt
"llms_txt_full_filename": "llms-full.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
manager.set_config(config)
# Create directories (but no _sources)
outdir = tmp_path / "build"
srcdir = tmp_path / "source"
outdir.mkdir()
srcdir.mkdir()
# Mock env with documents
class MockEnv:
all_docs = {"index": None}
titles = {"index": type("TitleNode", (), {"astext": lambda self: "Home"})()}
toctree_includes = {"index": []}
manager.set_env(MockEnv())
manager.set_master_doc("index")
manager.update_page_title("index", "Home")
# Capture warnings
captured_warnings = []
def capture_warning(message, *args, **kwargs):
if "_sources" in str(message):
captured_warnings.append(message)
with patch("sphinx_llms_txt.manager.logger.warning", side_effect=capture_warning):
# Call combine_sources
manager.combine_sources(str(outdir), str(srcdir))
# Verify NO warning was logged since llms_txt_full_file=False
assert (
len(captured_warnings) == 0
), "No warning should be logged when llms_txt_full_file=False"
# Verify llms.txt was still created
llms_txt = outdir / "llms.txt"
assert llms_txt.exists()
+175
View File
@@ -0,0 +1,175 @@
"""Test URI template functionality for llms.txt links."""
from sphinx_llms_txt import FileWriter
def test_uri_template_with_sources_dir(tmp_path):
"""Test that default template uses _sources links when sources_dir exists."""
build_dir = tmp_path / "build"
build_dir.mkdir()
# Create _sources directory to simulate its existence
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Mock app with html_sourcelink_suffix
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
config = Config()
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_uri_template": (
"{base_url}_sources/{docname}{suffix}{sourcelink_suffix}"
),
"html_baseurl": "https://example.com",
}
writer = FileWriter(config, str(build_dir), MockApp())
page_titles = {
"index": "Home Page",
"about": "About Us",
}
# Page order with suffixes (simulating _sources files exist)
page_order = [("index", ".rst"), ("about", ".md")]
writer.write_verbose_info_to_file(page_order, page_titles, 0, sources_dir)
# Check that the file was created
verbose_file = build_dir / "llms.txt"
assert verbose_file.exists()
# Read the file content
with open(verbose_file, "r", encoding="utf-8") as f:
content = f.read()
# Should link to _sources files
assert "- [Home Page](https://example.com/_sources/index.rst.txt)" in content
assert "- [About Us](https://example.com/_sources/about.md.txt)" in content
def test_uri_template_without_sources_dir(tmp_path):
"""
Test that HTML template is used when sources_dir doesn't exist and no custom
template.
"""
build_dir = tmp_path / "build"
build_dir.mkdir()
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
# No custom template set
"html_baseurl": "https://example.com",
}
writer = FileWriter(config, str(build_dir))
page_titles = {
"index": "Home Page",
"about": "About Us",
}
# Page order without suffixes (simulating no _sources)
page_order = [("index", None), ("about", None)]
# Pass None for sources_dir to simulate it doesn't exist
writer.write_verbose_info_to_file(page_order, page_titles, 0, None)
# Check that the file was created
verbose_file = build_dir / "llms.txt"
assert verbose_file.exists()
# Read the file content
with open(verbose_file, "r", encoding="utf-8") as f:
content = f.read()
# Should fallback to HTML links
assert "- [Home Page](https://example.com/index.html)" in content
assert "- [About Us](https://example.com/about.html)" in content
def test_uri_template_custom(tmp_path):
"""Test that custom URI template works correctly."""
build_dir = tmp_path / "build"
build_dir.mkdir()
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Mock app with html_sourcelink_suffix
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
config = Config()
# Custom template that uses different path
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_uri_template": "{base_url}raw/{docname}{suffix}",
"html_baseurl": "https://example.com/",
}
writer = FileWriter(config, str(build_dir), MockApp())
page_titles = {
"index": "Home Page",
}
page_order = [("index", ".rst")]
writer.write_verbose_info_to_file(page_order, page_titles, 0, sources_dir)
verbose_file = build_dir / "llms.txt"
with open(verbose_file, "r", encoding="utf-8") as f:
content = f.read()
# Should use custom template
assert "- [Home Page](https://example.com/raw/index.rst)" in content
def test_uri_template_invalid_fallback(tmp_path):
"""
Test that invalid template falls back to default sources template when
sources_dir exists.
"""
build_dir = tmp_path / "build"
build_dir.mkdir()
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Mock app with html_sourcelink_suffix
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
config = Config()
# Invalid template with typo in variable name
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_uri_template": "{base_urll}/{docname}",
"html_baseurl": "https://example.com",
}
writer = FileWriter(config, str(build_dir), MockApp())
page_titles = {
"index": "Home Page",
}
page_order = [("index", ".rst")]
writer.write_verbose_info_to_file(page_order, page_titles, 0, sources_dir)
verbose_file = build_dir / "llms.txt"
with open(verbose_file, "r", encoding="utf-8") as f:
content = f.read()
# Should fallback to default sources template
assert "- [Home Page](https://example.com/_sources/index.rst.txt)" in content