Compare commits

..
Author SHA1 Message Date
Jared Dillard 8625868df7 add soft transfer 2025-05-18 21:19:19 -07:00
Jared Dillard b2aa7897d3 Clean up 2025-05-18 21:08:58 -07:00
Jared Dillard 12d695015e change features to how it works 2025-05-18 21:04:59 -07:00
Jared Dillard 29c400e122 fix linter 2025-05-18 20:55:54 -07:00
Jared Dillard 834a57a158 add missing file 2025-05-18 20:54:55 -07:00
Jared Dillard efbe8e0cda revert dev env 2025-05-18 20:54:45 -07:00
Jared Dillard 68860c7dda clean up readme 2025-05-18 20:51:52 -07:00
Jared Dillard 055261b3bb fix linter 2025-05-18 20:50:38 -07:00
Jared Dillard a274621b32 fix linter 2025-05-18 20:49:44 -07:00
Jared Dillard 2f62461703 Add advanced configuration 2025-05-18 20:47:58 -07:00
Jared Dillard da48421b03 Clean up project name 2025-05-18 20:47:19 -07:00
Jared Dillard 9385670cfe Move readme content to index.rst 2025-05-18 20:46:51 -07:00
22 changed files with 262 additions and 2856 deletions
+5 -5
View File
@@ -10,9 +10,9 @@ jobs:
pre-commit: pre-commit:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v5 - uses: actions/checkout@v4
- name: Set up Python 3.10 - name: Set up Python 3.10
uses: actions/setup-python@v6 uses: actions/setup-python@v5
with: with:
python-version: "3.10" python-version: "3.10"
- uses: pre-commit/action@v3.0.1 - uses: pre-commit/action@v3.0.1
@@ -23,17 +23,17 @@ jobs:
python-version: ['3.9', '3.10', '3.11', '3.12'] python-version: ['3.9', '3.10', '3.11', '3.12']
steps: steps:
- uses: actions/checkout@v5 - uses: actions/checkout@v4
- name: Set up Python ${{ matrix.python-version }} - name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v6 uses: actions/setup-python@v5
with: with:
python-version: ${{ matrix.python-version }} python-version: ${{ matrix.python-version }}
- name: Install dependencies - name: Install dependencies
run: | run: |
python -m pip install --upgrade pip python -m pip install --upgrade pip
pip install -e . --group dev pip install -e ".[dev]"
# - name: Run mypy # - name: Run mypy
# run: | # run: |
-78
View File
@@ -1,84 +1,6 @@
Changelog Changelog
========= =========
0.6.0
-----
- Improve _sources directory handling
`#47 <https://github.com/jdillard/sphinx-llms-txt/pull/47>`_
0.5.3
-----
- Make sphinx a required dependency since there are imports from Sphinx
`#44 <https://github.com/jdillard/sphinx-llms-txt/pull/44>`_
0.5.2
-----
- Remove support for singlehtml
`#40 <https://github.com/jdillard/sphinx-llms-txt/pull/40>`_
0.5.1
-----
- Only allow builders that have _sources directory
`#38 <https://github.com/jdillard/sphinx-llms-txt/pull/38>`_
0.5.0
-----
- Add :ref:`block_level_ignore` and :ref:`page_level_ignore`
`#33 <https://github.com/jdillard/sphinx-llms-txt/pull/33>`_
- Add :confval:`llms_txt_full_size_policy` configuration option to control behavior when :confval:`llms_txt_full_max_size` is exceeded.
`#35 <https://github.com/jdillard/sphinx-llms-txt/pull/35>`_
0.4.1
-----
- Fix include paths and spacing
`#31 <https://github.com/jdillard/sphinx-llms-txt/pull/31>`_
0.4.0
-----
- Add support for including source code files with :confval:`llms_txt_code_files` and :confval:`llms_txt_code_base_path` configuration options
`#24 <https://github.com/jdillard/sphinx-llms-txt/pull/24>`_
0.3.2
-----
- Fix image paths to deployed images
`#30 <https://github.com/jdillard/sphinx-llms-txt/pull/30>`_
0.3.1
-----
- Fix issue when ``source_suffix`` equals ``source_link_suffix``
`#29 <https://github.com/jdillard/sphinx-llms-txt/pull/29>`_
0.3.0
-----
- Use first paragraph as default for ``llms_txt_summary``
`#22 <https://github.com/jdillard/sphinx-llms-txt/pull/22>`_
0.2.4
-----
- Support source file suffix detection
`#21 <https://github.com/jdillard/sphinx-llms-txt/pull/21>`_
0.2.3
-----
- Remove ``get_and_resolve_toctree`` method
`#19 <https://github.com/jdillard/sphinx-llms-txt/pull/19>`_
- Simplify ``_sources`` lookup
`#18 <https://github.com/jdillard/sphinx-llms-txt/pull/18>`_
- Add sphinx docs
`#16 <https://github.com/jdillard/sphinx-llms-txt/pull/16>`_
0.2.2 0.2.2
----- -----
+1 -7
View File
@@ -1,20 +1,14 @@
# Sphinx llms.txt generator # Sphinx llms.txt generator
A Sphinx extension that generates a summary `llms.txt` file and a single combined documentation `llms-full.txt` file. A Sphinx extension that generates a summary `llms.txt` file, written in Markdown, and a single combined documentation `llms-full.txt` file, written in reStructuredText.
[![PyPI version](https://img.shields.io/pypi/v/sphinx-llms-txt.svg)](https://pypi.python.org/pypi/sphinx-llms-txt) [![PyPI version](https://img.shields.io/pypi/v/sphinx-llms-txt.svg)](https://pypi.python.org/pypi/sphinx-llms-txt)
[![Conda Version](https://img.shields.io/conda/vn/conda-forge/sphinx-llms-txt.svg)](https://anaconda.org/conda-forge/sphinx-llms-txt)
[![Downloads](https://static.pepy.tech/badge/sphinx-llms-txt/month)](https://pepy.tech/project/sphinx-llms-txt) [![Downloads](https://static.pepy.tech/badge/sphinx-llms-txt/month)](https://pepy.tech/project/sphinx-llms-txt)
[![Parallel Safe](https://img.shields.io/badge/parallel%20safe-true-brightgreen)](#)
## Documentation ## Documentation
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions. See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
## Contributing
Pull Requests welcome! See [Contributing](https://sphinx-llms-txt.readthedocs.io/en/latest/contributing.html) for instructions on how best to contribute.
## License ## License
MIT License - see LICENSE file for details. MIT License - see LICENSE file for details.
+4 -136
View File
@@ -76,26 +76,15 @@ Handling Large Documentation
^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
For very large documentation sets, generating the full documentation file might exceed reasonable size limits. For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
You can set a maximum line count and control what happens when that limit is exceeded: You can set a maximum line count:
.. code-block:: python .. code-block:: python
llms_txt_full_max_size = 10000 # Maximum 10,000 lines llms_txt_full_max_size = 10000 # Maximum 10,000 lines
llms_txt_full_size_policy = "warn_skip" # Default behavior
The ``llms_txt_full_size_policy`` setting controls both the log level and action taken when the size limit is exceeded. If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
It uses the format ``"<loglevel>_<action>"``:
**Log levels:** .. tip:: Use :ref:`excluding_content` to remove less relevant pages.
- ``warn``: Log as a warning (default)
- ``info``: Log as informational message
**Actions:**
- ``skip``: Don't create the file (default)
- ``keep``: Create the file anyway, ignoring the size limit
- ``note``: Create a placeholder file explaining why the full file wasn't generated
.. tip:: Use :ref:`excluding_content` to remove less relevant pages and reduce the file size.
.. _custom_directive_handling: .. _custom_directive_handling:
@@ -124,13 +113,6 @@ This ensures that paths in your custom directives are properly resolved in the g
Excluding Content Excluding Content
^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^^
There are several ways to exclude content from the generated ``llms-full.txt`` file:
.. _global_exclusion:
Global Page Exclusion
~~~~~~~~~~~~~~~~~~~~~~
You can exclude specific pages from being included in the generated files: You can exclude specific pages from being included in the generated files:
.. code-block:: python .. code-block:: python
@@ -142,111 +124,6 @@ You can exclude specific pages from being included in the generated files:
] ]
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption. This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
It can also be used to reduce the size of llms-full.txt.
.. _page_level_ignore:
Page-Level Ignore Metadata
~~~~~~~~~~~~~~~~~~~~~~~~~~~
You can exclude individual pages by adding metadata at the top of any reStructuredText file:
.. code-block:: restructuredtext
:llms-txt-ignore: true
Page Title
==========
This entire page will be excluded from llms-full.txt
When this metadata is present, the entire page is skipped during processing.
.. _block_level_ignore:
Block-Level Ignore Directives
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
You can exclude specific sections within a page using ignore directives:
.. code-block:: restructuredtext
Page Title
==========
This content will be included in llms-full.txt.
.. llms-txt-ignore-start
This content will be excluded from llms-full.txt.
Section To Ignore
-----------------
This entire section and any nested content will be ignored.
.. code-block:: python
# This code block will also be ignored
def ignored_function():
pass
.. llms-txt-ignore-end
This content will be included again.
Block-level ignores can be useful for:
- Removing internal notes or TODOs
- Hiding implementation details while keeping user-facing documentation
.. note::
- Multiple ignore blocks can be used within the same file
- Ignore directives work with any indentation level
.. _including_code_files:
Including Source Code Files
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
You can include source code files from your project at the end of :confval:`llms_txt_full_filename`.
Use include/exclude syntax to precisely control which files are included:
.. code-block:: python
llms_txt_code_files = [
"+:src/**/*.py", # Include all Python files in src
"-:src/**/__pycache__/**", # Exclude Python cache files
]
Pattern syntax:
- **+:pattern**: Include files matching the pattern. Processed first to collect matching files.
- **-:pattern**: Exclude files matching the pattern. Applied to filter out unwanted files.
Code files are processed as follows:
- **Glob patterns**: Use standard glob patterns (``*``, ``**``, ``?``) to match files
- **Relative paths**: Patterns are resolved relative to your Sphinx source directory
- **Formatting**: Each file is presented with a title and syntax-highlighted code block
.. _customizing_code_paths:
Customizing Code File Paths
~~~~~~~~~~~~~~~~~~~~~~~~~~~~
By default, the extension automatically detects the relative path from your Sphinx source directory to the git root and strips that prefix from displayed file paths. You can customize this behavior:
.. code-block:: python
# Manually specify base path to strip
llms_txt_code_base_path = "../../"
# Disable path stripping entirely
llms_txt_code_base_path = ""
This helps create cleaner, more readable file paths in the generated documentation.
.. _using_html_baseurl: .. _using_html_baseurl:
@@ -269,7 +146,7 @@ Integration Examples
Complete Configuration Example Complete Configuration Example
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Here's a complete example showing multiple :doc:`configuration-values`: Here's a complete example showing multiple configuration options:
.. code-block:: python .. code-block:: python
@@ -277,7 +154,6 @@ Here's a complete example showing multiple :doc:`configuration-values`:
llms_txt_filename = "ai-summary.txt" llms_txt_filename = "ai-summary.txt"
llms_txt_full_filename = "ai-full-docs.txt" llms_txt_full_filename = "ai-full-docs.txt"
llms_txt_full_max_size = 50000 llms_txt_full_max_size = 50000
llms_txt_full_size_policy = "warn_note"
# Content customization # Content customization
llms_txt_title = "Project Documentation for AI Assistants" llms_txt_title = "Project Documentation for AI Assistants"
@@ -292,11 +168,3 @@ Here's a complete example showing multiple :doc:`configuration-values`:
# Content filtering # Content filtering
llms_txt_exclude = ["search", "genindex", "404", "private_*"] llms_txt_exclude = ["search", "genindex", "404", "private_*"]
# Source code inclusion with include/exclude patterns
llms_txt_code_files = [
"+:../../src/**/*.py", # Include Python files
"+:../../config/*.yaml", # Include config files
"-:../../src/**/__pycache__/**", # Exclude cache files
]
llms_txt_code_base_path = "../../"
+1 -6
View File
@@ -15,7 +15,6 @@ import subprocess
project = "sphinx-llms-txt" project = "sphinx-llms-txt"
copyright = "Jared Dillard" copyright = "Jared Dillard"
author = "Jared Dillard" author = "Jared Dillard"
llms_txt_code_files = ["+:../../sphinx_llms_txt/*.py"]
llms_txt_summary = """ llms_txt_summary = """
A Sphinx extension that generates a summary llms.txt file,written in Markdown, A Sphinx extension that generates a summary llms.txt file,written in Markdown,
and a single combined documentation llms-full.txt file, written in reStructuredText. and a single combined documentation llms-full.txt file, written in reStructuredText.
@@ -82,11 +81,7 @@ html_theme = "furo"
# further. For a list of options available for each theme, see the # further. For a list of options available for each theme, see the
# documentation. # documentation.
# #
html_theme_options = { html_theme_options = {}
"source_repository": "https://github.com/jdillard/sphinx-llms-txt/",
"source_branch": "main",
"source_directory": "docs/source/",
}
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/" html_baseurl = "https://sphinx-llms-txt.readthedocs.org/"
+4 -34
View File
@@ -24,22 +24,11 @@ Project Configuration Values
- **Type**: integer or ``None`` - **Type**: integer or ``None``
- **Default**: ``None`` (no limit) - **Default**: ``None`` (no limit)
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``. - **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
Behavior when exceeded is controlled by :confval:`llms_txt_full_size_policy`. If exceeded, the file is skipped and a warning is shown, but the build still completes.
See :ref:`handling_large_documentation`. See :ref:`handling_large_documentation`.
.. versionadded:: 0.2.0 .. versionadded:: 0.2.0
.. confval:: llms_txt_full_size_policy
- **Type**: string
- **Default**: ``'warn_skip'``
- **Description**: Controls what happens when :confval:`llms_txt_full_max_size` is exceeded.
Format is ``<loglevel>_<action>``. Log levels: ``warn``, ``info``.
Actions: ``skip``, ``keep``, ``note``.
See :ref:`handling_large_documentation`.
.. versionadded:: 0.5.0
.. confval:: llms_txt_file .. confval:: llms_txt_file
- **Type**: boolean - **Type**: boolean
@@ -78,8 +67,8 @@ Project Configuration Values
.. confval:: llms_txt_summary .. confval:: llms_txt_summary
- **Type**: string - **Type**: string or ``None``
- **Default**: The first paragraph in the root document, else an empty string - **Default**: ``None``
- **Description**: Optional, but recommended, summary description for ``llms.txt``. - **Description**: Optional, but recommended, summary description for ``llms.txt``.
See :ref:`custom_summary`. See :ref:`custom_summary`.
@@ -89,26 +78,7 @@ Project Configuration Values
- **Type**: list of strings - **Type**: list of strings
- **Default**: ``[]`` - **Default**: ``[]``
- **Description**: A list of pages to ignore using glob patterns. - **Description**: A list of pages to ignore.
See :ref:`excluding_content`. See :ref:`excluding_content`.
.. versionadded:: 0.2.1 .. versionadded:: 0.2.1
.. confval:: llms_txt_code_files
- **Type**: list of strings
- **Default**: ``[]``
- **Description**: A list of glob patterns that appends source code files to :confval:`llms_txt_full_filename`.
See :ref:`including_code_files`.
.. versionadded:: 0.4.0
.. confval:: llms_txt_code_base_path
- **Type**: string or ``None``
- **Default**: ``None`` (auto-detect from git root)
- **Description**: Base path to strip from code file paths when displaying titles.
When ``None``, automatically detects the relative path from the Sphinx source
directory to the git root and strips that prefix from file paths.
.. versionadded:: 0.4.0
+1 -1
View File
@@ -19,7 +19,7 @@ Local development
.. code-block:: console .. code-block:: console
pip install -e . --group dev pip install -e ".[dev]"
#. Install pre-commit Git hook scripts: #. Install pre-commit Git hook scripts:
+25 -13
View File
@@ -1,6 +1,11 @@
Getting Started Getting Started
=============== ===============
Demo
----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
Installation Installation
------------ ------------
@@ -10,12 +15,6 @@ Directly install via ``pip`` by using:
pip install sphinx-llms-txt pip install sphinx-llms-txt
Or with ``conda`` via ``conda-forge``:
.. code::
conda install -c conda-forge sphinx-llms-txt
Usage Usage
----- -----
@@ -27,12 +26,25 @@ Add the extension to your Sphinx configuration (``conf.py``):
'sphinx_llms_txt', 'sphinx_llms_txt',
] ]
After the HTML finishes building, **sphinx-llms-txt** will output the location of the output files:: Once added, the extension will automatically generate the LLMs.txt files during the build process.
sphinx-llms-txt: Created /path/to/_build/html/llms-full.txt with 45 sources and 6879 lines
sphinx-llms-txt: created /path/to/_build/html/llms.txt
.. tip:: Make sure to confirm the accuracy of the output files after installs and upgrades.
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**. See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
How It Works
-----------
During the Sphinx build process:
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
3. **Path Resolution**: Transforms relative paths in directives to full paths
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
+1 -34
View File
@@ -3,26 +3,7 @@ Sphinx llms.txt Generator
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText. A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars| |PyPI version|
Demo
----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
Highlights
----------
1. **Content Collection**: Quickly gathers content from _sources, without needing a separate build
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
3. **Path Resolution**: Transforms relative paths in directives to full paths
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
5. **Content Filtering**: Allows you to exclude specific pages or sections
6. **Source Code**: Allows you to include specific source code files
.. toctree:: .. toctree::
:maxdepth: 2 :maxdepth: 2
@@ -34,22 +15,8 @@ Highlights
changelog changelog
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
.. _Sphinx: http://sphinx-doc.org/ .. _Sphinx: http://sphinx-doc.org/
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg .. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
:target: https://pypi.python.org/pypi/sphinx-llms-txt :target: https://pypi.python.org/pypi/sphinx-llms-txt
:alt: Latest PyPi Version :alt: Latest PyPi Version
.. |Conda Version| image:: https://img.shields.io/conda/vn/conda-forge/sphinx-llms-txt.svg
:target: https://anaconda.org/conda-forge/sphinx-llms-txt
:alt: Latest Conda Version
.. |Downloads| image:: https://static.pepy.tech/badge/sphinx-llms-txt/month
:target: https://pepy.tech/project/sphinx-llms-txt
:alt: PyPi Downloads per month
.. |Parallel Safe| image:: https://img.shields.io/badge/parallel%20safe-true-brightgreen
:target: #
:alt: Parallel read/write safe
.. |GitHub Stars| image:: https://img.shields.io/github/stars/jdillard/sphinx-llms-txt?style=social
:target: https://github.com/jdillard/sphinx-llms-txt
:alt: GitHub Repository stars
+2 -4
View File
@@ -26,16 +26,13 @@ classifiers = [
license = {text = "MIT"} license = {text = "MIT"}
readme = "README.md" readme = "README.md"
dynamic = ["version"] dynamic = ["version"]
dependencies = [
"sphinx",
]
[project.urls] [project.urls]
download = "https://pypi.org/project/sphinx-llms-txt/" download = "https://pypi.org/project/sphinx-llms-txt/"
source = "https://github.com/jdillard/sphinx-llms-txt" source = "https://github.com/jdillard/sphinx-llms-txt"
changelog = "https://github.com/jdillard/sphinx-llms-txt/blob/master/CHANGELOG.rst" changelog = "https://github.com/jdillard/sphinx-llms-txt/blob/master/CHANGELOG.rst"
[dependency-groups] [project.optional-dependencies]
dev = [ dev = [
"pytest>=7.0.0", "pytest>=7.0.0",
"black", "black",
@@ -43,6 +40,7 @@ dev = [
"mypy", "mypy",
"isort", "isort",
"pre-commit", "pre-commit",
"sphinx",
] ]
test = [ test = [
"pytest>=7.0.0", "pytest>=7.0.0",
+8 -54
View File
@@ -1,14 +1,5 @@
""" """
Sphinx extension that generates llms.txt and llms-full.txt files for LLM consumption. Sphinx extension to create a combined sources file (llms-full.txt)
This extension collects documentation content from Sphinx projects and generates
two output files:
- llms.txt: A concise Markdown summary with project overview and page links
- llms-full.txt: A comprehensive reStructuredText file containing all documentation
content with resolved includes and path references
The extension processes content during the build phase, handles page-level and
block-level ignore directives, and can optionally include source code files.
""" """
from typing import Any, Dict from typing import Any, Dict
@@ -21,7 +12,7 @@ from .manager import LLMSFullManager
from .processor import DocumentProcessor from .processor import DocumentProcessor
from .writer import FileWriter from .writer import FileWriter
__version__ = "0.6.0" __version__ = "0.2.2"
# Export classes needed by tests # Export classes needed by tests
__all__ = [ __all__ = [
@@ -34,21 +25,9 @@ __all__ = [
# Global manager instance # Global manager instance
_manager = LLMSFullManager() _manager = LLMSFullManager()
# Store root document first paragraph
_root_first_paragraph = ""
def doctree_resolved(app: Sphinx, doctree, docname: str): def doctree_resolved(app: Sphinx, doctree, docname: str):
"""Called when a docname has been resolved to a document.""" """Called when a docname has been resolved to a document."""
global _root_first_paragraph
# Check for llms-txt-ignore metadata at the page level
if hasattr(app.env, "metadata") and docname in app.env.metadata:
metadata = app.env.metadata[docname]
if metadata.get("llms-txt-ignore", "").lower() in ("true", "1", "yes"):
_manager.mark_page_ignored(docname)
return
# Extract title from the document # Extract title from the document
title = None title = None
# findall() returns a generator, convert to list to check if it has elements # findall() returns a generator, convert to list to check if it has elements
@@ -59,14 +38,6 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
if title: if title:
_manager.update_page_title(docname, title) _manager.update_page_title(docname, title)
# Extract first paragraph from root document
if docname == app.config.master_doc:
for node in doctree.traverse(nodes.paragraph):
first_para = node.astext()
if first_para:
_root_first_paragraph = first_para
break
def build_finished(app: Sphinx, exception): def build_finished(app: Sphinx, exception):
"""Called when the build is finished.""" """Called when the build is finished."""
@@ -76,25 +47,17 @@ def build_finished(app: Sphinx, exception):
_manager.set_master_doc(app.config.master_doc) _manager.set_master_doc(app.config.master_doc)
_manager.set_app(app) _manager.set_app(app)
# Get the summary - use configured value or extracted first paragraph
summary = app.config.llms_txt_summary
if summary is None:
summary = _root_first_paragraph
# Set up configuration # Set up configuration
config = { config = {
"llms_txt_file": app.config.llms_txt_file, "llms_txt_file": app.config.llms_txt_file,
"llms_txt_filename": app.config.llms_txt_filename, "llms_txt_filename": app.config.llms_txt_filename,
"llms_txt_title": app.config.llms_txt_title, "llms_txt_title": app.config.llms_txt_title,
"llms_txt_summary": summary, "llms_txt_summary": app.config.llms_txt_summary,
"llms_txt_full_file": app.config.llms_txt_full_file, "llms_txt_full_file": app.config.llms_txt_full_file,
"llms_txt_full_filename": app.config.llms_txt_full_filename, "llms_txt_full_filename": app.config.llms_txt_full_filename,
"llms_txt_full_max_size": app.config.llms_txt_full_max_size, "llms_txt_full_max_size": app.config.llms_txt_full_max_size,
"llms_txt_full_size_policy": app.config.llms_txt_full_size_policy,
"llms_txt_directives": app.config.llms_txt_directives, "llms_txt_directives": app.config.llms_txt_directives,
"llms_txt_exclude": app.config.llms_txt_exclude, "llms_txt_exclude": app.config.llms_txt_exclude,
"llms_txt_code_files": app.config.llms_txt_code_files,
"llms_txt_code_base_path": app.config.llms_txt_code_base_path,
"html_baseurl": getattr(app.config, "html_baseurl", ""), "html_baseurl": getattr(app.config, "html_baseurl", ""),
} }
_manager.set_config(config) _manager.set_config(config)
@@ -113,33 +76,24 @@ def build_finished(app: Sphinx, exception):
def setup(app: Sphinx) -> Dict[str, Any]: def setup(app: Sphinx) -> Dict[str, Any]:
"""Set up the Sphinx extension.""" """Set up the Sphinx extension."""
# Add configuration options
app.add_config_value("llms_txt_file", True, "env") app.add_config_value("llms_txt_file", True, "env")
app.add_config_value("llms_txt_filename", "llms.txt", "env") app.add_config_value("llms_txt_filename", "llms.txt", "env")
app.add_config_value("llms_txt_full_file", True, "env") app.add_config_value("llms_txt_full_file", True, "env")
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env") app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
app.add_config_value("llms_txt_full_max_size", None, "env") app.add_config_value("llms_txt_full_max_size", None, "env")
app.add_config_value("llms_txt_full_size_policy", "warn_skip", "env")
app.add_config_value("llms_txt_directives", [], "env") app.add_config_value("llms_txt_directives", [], "env")
app.add_config_value("llms_txt_title", None, "env") app.add_config_value("llms_txt_title", None, "env")
app.add_config_value("llms_txt_summary", None, "env") app.add_config_value("llms_txt_summary", None, "env")
app.add_config_value("llms_txt_exclude", [], "env") app.add_config_value("llms_txt_exclude", [], "env")
app.add_config_value("llms_txt_code_files", [], "env")
app.add_config_value("llms_txt_code_base_path", None, "env")
def builder_inited(app):
"""Used to limit what builders are allowed to run the extension."""
allowed_builders = ["html", "dirhtml"]
if hasattr(app, "builder") and app.builder.name in allowed_builders:
# Reset manager and root paragraph for each build
global _manager, _root_first_paragraph
_manager = LLMSFullManager()
_root_first_paragraph = ""
# Connect to Sphinx events
app.connect("doctree-resolved", doctree_resolved) app.connect("doctree-resolved", doctree_resolved)
app.connect("build-finished", build_finished) app.connect("build-finished", build_finished)
app.connect("builder-inited", builder_inited) # Reset manager for each build
global _manager
_manager = LLMSFullManager()
return { return {
"version": __version__, "version": __version__,
+24 -123
View File
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
""" """
import fnmatch import fnmatch
from typing import Any, Dict, List, Tuple from typing import Any, Dict, List
from sphinx.environment import BuildEnvironment from sphinx.environment import BuildEnvironment
from sphinx.util import logging from sphinx.util import logging
@@ -19,7 +19,6 @@ class DocumentCollector:
self.master_doc: str = None self.master_doc: str = None
self.env: BuildEnvironment = None self.env: BuildEnvironment = None
self.config: Dict[str, Any] = {} self.config: Dict[str, Any] = {}
self.app = None
def set_master_doc(self, master_doc: str): def set_master_doc(self, master_doc: str):
"""Set the master document name.""" """Set the master document name."""
@@ -38,79 +37,8 @@ class DocumentCollector:
"""Set configuration options.""" """Set configuration options."""
self.config = config self.config = config
def set_app(self, app): def get_page_order(self) -> List[str]:
"""Set the Sphinx application reference.""" """Get the correct page order from the toctree structure."""
self.app = app
def _get_source_suffixes(self):
"""Get all valid source file suffixes from Sphinx configuration.
Returns:
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
"""
if not self.app:
return [".rst"] # Default fallback
source_suffix = self.app.config.source_suffix
if isinstance(source_suffix, dict):
return list(source_suffix.keys())
elif isinstance(source_suffix, list):
return source_suffix
else:
return [source_suffix] # String format
def _get_docname_suffix(self, docname: str, sources_dir) -> str:
"""
Determine the source suffix for a given docname by checking which
file exists.
Args:
docname: The document name to check
sources_dir: Path to the _sources directory
Returns:
The source suffix if found, or None if no matching file exists
"""
if not sources_dir or not sources_dir.exists():
return None
# Get the source link suffix from Sphinx config
source_link_suffix = ""
if self.app and hasattr(self.app.config, "html_sourcelink_suffix"):
source_link_suffix = self.app.config.html_sourcelink_suffix
# Handle empty string case specially
if source_link_suffix == "":
source_link_suffix = "" # Keep it empty
elif not source_link_suffix.startswith("."):
source_link_suffix = "." + source_link_suffix
# Get the source file suffixes from Sphinx config
source_suffixes = self._get_source_suffixes()
# Try to find the source file with any of the valid source suffixes
for src_suffix in source_suffixes:
# Avoid duplicate extensions when source_suffix == source_link_suffix
if src_suffix == source_link_suffix:
candidate_file = sources_dir / f"{docname}{src_suffix}"
else:
candidate_file = (
sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
)
if candidate_file.exists():
return src_suffix
return None
def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
"""Get the correct page order from the toctree structure.
Args:
sources_dir: Optional path to _sources directory for suffix detection
Returns:
List of tuples (docname, source_suffix) in toctree order
"""
if not self.env or not self.master_doc: if not self.env or not self.master_doc:
return [] return []
@@ -124,12 +52,9 @@ class DocumentCollector:
visited.add(docname) visited.add(docname)
# Add the current document with its suffix # Add the current document
if docname not in [doc for doc, _ in page_order]: if docname not in page_order:
suffix = None page_order.append(docname)
if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix))
# Check for toctree entries in this document # Check for toctree entries in this document
try: try:
@@ -140,33 +65,20 @@ class DocumentCollector:
): ):
for child_docname in self.env.toctree_includes[docname]: for child_docname in self.env.toctree_includes[docname]:
collect_from_toctree(child_docname) collect_from_toctree(child_docname)
# Try to use dependencies to find related documents else:
elif ( # Fallback: try to resolve and parse the toctree
hasattr(self.env, "dependencies") toctree = self.env.get_and_resolve_toctree(docname, None)
and docname in self.env.dependencies if toctree:
): from docutils import nodes
# Extract the dependent documents from the dependencies dict
for child_docname in self.env.dependencies[docname]:
# Only add documents actually in the document set
if (
hasattr(self.env, "all_docs")
and child_docname in self.env.all_docs
):
collect_from_toctree(child_docname)
# Fallback to titles or other available references
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
# Get all document names
all_docnames = list(self.env.all_docs.keys())
# Look for documents that might be related (have similar paths) for node in list(toctree.findall(nodes.reference)):
current_prefix = "/".join(docname.split("/")[:-1]) if "refuri" in node.attributes:
if current_prefix: refuri = node.attributes["refuri"]
for child_docname in all_docnames: if refuri and refuri.endswith(".html"):
# Documents in the same directory might be related child_docname = refuri[:-5] # Remove .html
if ( if (
child_docname.startswith(current_prefix) child_docname != docname
and child_docname != docname ): # Avoid circular references
):
collect_from_toctree(child_docname) collect_from_toctree(child_docname)
except Exception as e: except Exception as e:
logger.debug(f"Could not get toctree for {docname}: {e}") logger.debug(f"Could not get toctree for {docname}: {e}")
@@ -176,33 +88,22 @@ class DocumentCollector:
# Add any remaining documents not in the toctree (sorted) # Add any remaining documents not in the toctree (sorted)
if hasattr(self.env, "all_docs"): if hasattr(self.env, "all_docs"):
processed_docnames = {doc for doc, _ in page_order}
remaining = sorted( remaining = sorted(
[ [doc for doc in self.env.all_docs.keys() if doc not in page_order]
doc
for doc in self.env.all_docs.keys()
if doc not in processed_docnames
]
) )
for docname in remaining: page_order.extend(remaining)
suffix = None
if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix))
return page_order return page_order
def filter_excluded_pages( def filter_excluded_pages(self, page_order: List[str]) -> List[str]:
self, page_order: List[Tuple[str, str]]
) -> List[Tuple[str, str]]:
"""Filter out excluded pages from the page order.""" """Filter out excluded pages from the page order."""
exclude_patterns = self.config.get("llms_txt_exclude") exclude_patterns = self.config.get("llms_txt_exclude")
if exclude_patterns: if exclude_patterns:
return [ return [
(docname, suffix) page
for docname, suffix in page_order for page in page_order
if not any( if not any(
self._match_exclude_pattern(docname, pattern) self._match_exclude_pattern(page, pattern)
for pattern in exclude_patterns for pattern in exclude_patterns
) )
] ]
+112 -780
View File
File diff suppressed because it is too large Load Diff
+9 -111
View File
@@ -44,10 +44,7 @@ class DocumentProcessor:
Returns: Returns:
Processed content with directives properly resolved Processed content with directives properly resolved
""" """
# First process llms-txt-ignore blocks # First process include directives
content = self._process_ignore_blocks(content)
# Then process include directives
content = self._process_includes(content, source_path) content = self._process_includes(content, source_path)
# Then process path directives (image, figure, etc.) # Then process path directives (image, figure, etc.)
@@ -97,14 +94,8 @@ class DocumentProcessor:
if not base_url: if not base_url:
return path return path
# Ensure base URL ends with slash
if not base_url.endswith("/"): if not base_url.endswith("/"):
base_url += "/" base_url += "/"
# Remove leading slash from path to avoid double slashes
if path.startswith("/"):
path = path[1:]
return f"{base_url}{path}" return f"{base_url}{path}"
def _is_absolute_or_url(self, path: str) -> bool: def _is_absolute_or_url(self, path: str) -> bool:
@@ -146,41 +137,8 @@ class DocumentProcessor:
prefix = match.group(1) # The entire directive prefix including whitespace prefix = match.group(1) # The entire directive prefix including whitespace
path = match.group(3).strip() # The path argument path = match.group(3).strip() # The path argument
# Handle URLs and data URIs - leave unchanged # Only process relative paths, not absolute paths or URLs
if path.startswith(("http://", "https://", "data:")): if not self._is_absolute_or_url(path):
return match.group(0)
# For ALL paths, check if image exists in _images first
# Extract filename from the path
filename = os.path.basename(path)
# Check if image exists in _images directory
# First determine the build directory from source_path
build_dir = None
if "_sources" in str(source_path):
# Extract build directory (parent of _sources)
path_parts = str(source_path).split("_sources/")
if len(path_parts) > 1:
build_dir = path_parts[0].rstrip("/")
# If we can determine the build directory, check if image exists in _images
if build_dir:
images_path = os.path.join(build_dir, "_images", filename)
if os.path.exists(images_path):
# Image exists in _images, use _images path
full_path = f"/_images/{filename}"
# Add base URL if configured
full_path = self._add_base_url(full_path, base_url)
return f"{prefix}{full_path}"
# Image doesn't exist in _images, handle based on path type
# Handle absolute paths (starting with /) - add base URL if configured
if path.startswith("/"):
# Add base URL to absolute paths if configured
full_path = self._add_base_url(path, base_url)
return f"{prefix}{full_path}"
# Handle relative paths with original logic for backward compatibility
# Special case for test files # Special case for test files
if is_test: if is_test:
# Add subdir/ prefix to match test expectations # Add subdir/ prefix to match test expectations
@@ -210,7 +168,9 @@ class DocumentProcessor:
elif rel_doc_dir: elif rel_doc_dir:
# Join with the original path to form full path relative # Join with the original path to form full path relative
# to srcdir # to srcdir
full_path = os.path.normpath(os.path.join(rel_doc_dir, path)) full_path = os.path.normpath(
os.path.join(rel_doc_dir, path)
)
else: else:
full_path = path full_path = path
@@ -220,12 +180,7 @@ class DocumentProcessor:
# Return the updated directive with the full path # Return the updated directive with the full path
return f"{prefix}{full_path}" return f"{prefix}{full_path}"
# Fallback for relative paths - add base URL if configured # If we couldn't resolve the path or it's already absolute, return unchanged
else:
full_path = self._add_base_url(path, base_url)
return f"{prefix}{full_path}"
# If we couldn't resolve the path, return unchanged
return match.group(0) return match.group(0)
# Replace directive paths in the content # Replace directive paths in the content
@@ -246,12 +201,9 @@ class DocumentProcessor:
""" """
possible_paths = [] possible_paths = []
# If it's an absolute path, treat it as relative to srcdir # If it's an absolute path, use it directly
if os.path.isabs(include_path): if os.path.isabs(include_path):
# Remove the leading slash and treat as relative to srcdir possible_paths.append(Path(include_path))
relative_path = include_path.lstrip("/")
if self.srcdir:
possible_paths.append((Path(self.srcdir) / relative_path).resolve())
else: else:
# Relative to the source file (in _sources directory) # Relative to the source file (in _sources directory)
possible_paths.append((source_path.parent / include_path).resolve()) possible_paths.append((source_path.parent / include_path).resolve())
@@ -292,9 +244,6 @@ class DocumentProcessor:
# Function to replace each include with content # Function to replace each include with content
def replace_include(match): def replace_include(match):
include_path = match.group(3) include_path = match.group(3)
directive_part = match.group(
1
) # The ".. include:: " part with leading whitespace
# Get all possible paths to try # Get all possible paths to try
possible_paths = self._resolve_include_paths(include_path, source_path) possible_paths = self._resolve_include_paths(include_path, source_path)
@@ -305,18 +254,7 @@ class DocumentProcessor:
if path_to_try.exists(): if path_to_try.exists():
with open(path_to_try, "r", encoding="utf-8") as f: with open(path_to_try, "r", encoding="utf-8") as f:
included_content = f.read() included_content = f.read()
# Find where the actual directive starts, after any whitespace
directive_start = directive_part.find("..")
if directive_start > 0:
# There's leading whitespace/newlines before the directive
leading_part = directive_part[:directive_start]
# Replace directive with content, preserving the structure
return leading_part + included_content
else:
# No leading whitespace, just return the content
return included_content return included_content
except Exception as e: except Exception as e:
logger.error( logger.error(
f"sphinx-llms-txt: Error reading include file {path_to_try}:" f"sphinx-llms-txt: Error reading include file {path_to_try}:"
@@ -328,48 +266,8 @@ class DocumentProcessor:
paths_tried = ", ".join(str(p) for p in possible_paths) paths_tried = ", ".join(str(p) for p in possible_paths)
logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}") logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}")
logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}") logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}")
# Preserve spacing structure for error message too
directive_start = match.group(1).find("..")
if directive_start > 0:
leading_part = match.group(1)[:directive_start]
return leading_part + f"[Include file not found: {include_path}]"
else:
return f"[Include file not found: {include_path}]" return f"[Include file not found: {include_path}]"
# Replace all includes with their content # Replace all includes with their content
processed_content = include_pattern.sub(replace_include, content) processed_content = include_pattern.sub(replace_include, content)
return processed_content return processed_content
def _process_ignore_blocks(self, content: str) -> str:
"""Process llms-txt-ignore-start/end blocks by removing their content.
Args:
content: The source content to process
Returns:
Processed content with ignore blocks removed
"""
# Process ignore blocks iteratively to handle nested cases correctly
while True:
# Pattern to match ignore blocks - handles whitespace and indentation
ignore_pattern = re.compile(
r"^\s*\.\.\s+llms-txt-ignore-start\s*\n" # Start directive line
r"(.*?)" # Content to ignore (non-greedy)
r"^\s*\.\.\s+llms-txt-ignore-end\s*$", # End directive line
re.MULTILINE | re.DOTALL,
)
# Find and remove one ignore block at a time
match = ignore_pattern.search(content)
if not match:
break
# Remove the matched block
content = content[: match.start()] + content[match.end() :]
# Clean up any extra blank lines that might be left
# Replace multiple consecutive newlines with at most 2 newlines
processed_content = re.sub(r"\n\n\n+", "\n\n", content)
return processed_content
+6 -13
View File
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
""" """
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Tuple, Union from typing import Any, Dict, List
from sphinx.application import Sphinx from sphinx.application import Sphinx
from sphinx.util import logging from sphinx.util import logging
@@ -37,24 +37,24 @@ class FileWriter:
f.write("\n".join(content_parts)) f.write("\n".join(content_parts))
logger.info( logger.info(
f"sphinx-llms-txt: Created {output_path} with {len(content_parts)}" f"sphinx-llms-txt: created {output_path} with {len(content_parts)}"
f" sources and {total_line_count} lines" f" sources and {total_line_count} lines"
) )
return True return True
except Exception as e: except Exception as e:
logger.error(f"sphinx-llms-txt: Error writing combined sources file: {e}") logger.error(f"sphinx-llm-txt: Error writing combined sources file: {e}")
return False return False
def write_verbose_info_to_file( def write_verbose_info_to_file(
self, self,
page_order: Union[List[str], List[Tuple[str, str]]], page_order: List[str],
page_titles: Dict[str, str], page_titles: Dict[str, str],
total_line_count: int = 0, total_line_count: int = 0,
) -> bool: ) -> bool:
"""Write summary information to the llms.txt file. """Write summary information to the llms.txt file.
Args: Args:
page_order: Ordered list of document names or (docname, suffix) tuples page_order: Ordered list of document names
page_titles: Dictionary mapping docnames to titles page_titles: Dictionary mapping docnames to titles
total_line_count: Total number of lines in the combined content total_line_count: Total number of lines in the combined content
@@ -88,8 +88,6 @@ class FileWriter:
if description: if description:
# Trim leading and trailing whitespace # Trim leading and trailing whitespace
description = description.strip() description = description.strip()
if description:
# Only add blockquote if description is not empty
# Replace newlines with newline + blockquote marker to maintain # Replace newlines with newline + blockquote marker to maintain
# blockquote formatting # blockquote formatting
description = description.replace("\n", "\n> ") description = description.replace("\n", "\n> ")
@@ -102,12 +100,7 @@ class FileWriter:
if not base_url.endswith("/"): if not base_url.endswith("/"):
base_url += "/" base_url += "/"
for item in page_order: for docname in page_order:
# Handle both old format (str) and new format (tuple)
if isinstance(item, tuple):
docname, _ = item
else:
docname = item
title = page_titles.get(docname, docname) title = page_titles.get(docname, docname)
f.write(f"- [{title}]({base_url}{docname}.html)\n") f.write(f"- [{title}]({base_url}{docname}.html)\n")
-2
View File
@@ -8,8 +8,6 @@ Welcome to Test Project's documentation!
page1 page1
page2 page2
page_with_include page_with_include
page_ignored_metadata
page_with_ignore_blocks
Indices and tables Indices and tables
================== ==================
@@ -1,16 +0,0 @@
:llms-txt-ignore: true
Page Ignored by Metadata
========================
This page should not appear in llms-full.txt because of the metadata directive.
Section 1
---------
This content should be completely ignored.
Section 2
---------
This content should also be ignored.
@@ -1,39 +0,0 @@
Page With Ignore Blocks
=======================
This content should appear in llms-full.txt.
.. llms-txt-ignore-start
This content should be ignored and not appear in llms-full.txt.
Section Ignored
---------------
This section should also be ignored.
.. llms-txt-ignore-end
This content after the ignore block should appear in llms-full.txt.
Another Section
---------------
This content should definitely appear.
.. llms-txt-ignore-start
Another ignored block with multiple lines.
- Item 1 (ignored)
- Item 2 (ignored)
.. code-block:: python
# This code should be ignored
def ignored_function():
pass
.. llms-txt-ignore-end
Final content that should appear.
-220
View File
@@ -1,220 +0,0 @@
"""Tests for llms-txt ignore features."""
from pathlib import Path
from sphinx_llms_txt import DocumentProcessor
def test_process_ignore_blocks():
"""Test that ignore blocks are properly removed from content."""
processor = DocumentProcessor({}, None)
content = """This content should remain.
.. llms-txt-ignore-start
This content should be removed.
Section Ignored
---------------
This section should also be removed.
.. llms-txt-ignore-end
This content should remain after the ignore block.
.. llms-txt-ignore-start
Another ignored block.
Multiple lines here.
.. llms-txt-ignore-end
Final content that should remain."""
processed = processor._process_ignore_blocks(content)
# Check that ignored content is removed
assert "This content should be removed." not in processed
assert "Section Ignored" not in processed
assert "Another ignored block." not in processed
assert "Multiple lines here." not in processed
# Check that non-ignored content remains
assert "This content should remain." in processed
assert "This content should remain after the ignore block." in processed
assert "Final content that should remain." in processed
def test_process_ignore_blocks_with_indentation():
"""Test that ignore blocks work with different indentation levels."""
processor = DocumentProcessor({}, None)
content = """Section Title
=============
Normal content.
.. llms-txt-ignore-start
Indented ignored content.
More indented content.
.. llms-txt-ignore-end
Back to normal content."""
processed = processor._process_ignore_blocks(content)
# Check that ignored content is removed
assert "Indented ignored content." not in processed
assert "More indented content." not in processed
# Check that non-ignored content remains
assert "Section Title" in processed
assert "Normal content." in processed
assert "Back to normal content." in processed
def test_process_ignore_blocks_multiple():
"""Test that multiple ignore blocks are handled correctly."""
processor = DocumentProcessor({}, None)
content = """Start content.
.. llms-txt-ignore-start
First ignore block.
.. llms-txt-ignore-end
Middle content that should remain.
.. llms-txt-ignore-start
Second ignore block.
.. llms-txt-ignore-end
End content."""
processed = processor._process_ignore_blocks(content)
# Check that ignored content is removed
assert "First ignore block." not in processed
assert "Second ignore block." not in processed
# Check that non-ignored content remains
assert "Start content." in processed
assert "Middle content that should remain." in processed
assert "End content." in processed
def test_build_with_ignore_features(basic_sphinx_app):
"""Test building HTML documentation with ignore features."""
app = basic_sphinx_app
app.build()
# Check if the output file was created
output_file = Path(app.outdir) / "test-llms-full.txt"
assert output_file.exists(), f"Output file {output_file} does not exist"
# Read the content of the output file
content = output_file.read_text()
# Check that page with metadata ignore is completely excluded
assert "Page Ignored by Metadata" not in content
assert "This page should not appear in llms-full.txt" not in content
# Check that page with ignore blocks has the right content
assert "Page With Ignore Blocks" in content
assert "This content should appear in llms-full.txt." in content
assert "This content after the ignore block should appear" in content
assert "Another Section" in content
assert "Final content that should appear." in content
# Check that ignored block content is not present
assert "This content should be ignored and not appear" not in content
assert "Section Ignored" not in content
assert "Another ignored block with multiple lines." not in content
assert "Item 1 (ignored)" not in content
assert "def ignored_function():" not in content
def test_manager_mark_page_ignored():
"""Test that manager can mark pages as ignored."""
from sphinx_llms_txt import LLMSFullManager
manager = LLMSFullManager()
# Initially no pages are ignored
assert len(manager.ignored_pages) == 0
# Mark a page as ignored
manager.mark_page_ignored("test_page")
# Check that page is in ignored set
assert "test_page" in manager.ignored_pages
assert len(manager.ignored_pages) == 1
# Mark another page as ignored
manager.mark_page_ignored("another_page")
# Check both pages are ignored
assert "test_page" in manager.ignored_pages
assert "another_page" in manager.ignored_pages
assert len(manager.ignored_pages) == 2
def test_process_ignore_blocks_empty_blocks():
"""Test that empty ignore blocks are handled correctly."""
processor = DocumentProcessor({}, None)
content = """Content before.
.. llms-txt-ignore-start
.. llms-txt-ignore-end
Content after."""
processed = processor._process_ignore_blocks(content)
# Check that content remains
assert "Content before." in processed
assert "Content after." in processed
# Check that we don't have excessive newlines
lines = processed.strip().split("\n")
non_empty_lines = [line for line in lines if line.strip()]
assert len(non_empty_lines) == 2
def test_ignore_metadata_affects_both_files(basic_sphinx_app):
"""Test that :llms-txt-ignore: true affects both files."""
app = basic_sphinx_app
# Enable both llms.txt and llms-full.txt file generation
app.config.llms_txt_file = True
app.config.llms_txt_filename = "test-llms.txt"
app.build()
# Check if both output files were created
llms_full_file = Path(app.outdir) / "test-llms-full.txt"
llms_summary_file = Path(app.outdir) / "test-llms.txt"
assert llms_full_file.exists(), f"Output file {llms_full_file} does not exist"
assert llms_summary_file.exists(), f"Output file {llms_summary_file} does not exist"
# Read the content of both files
llms_full_content = llms_full_file.read_text()
llms_summary_content = llms_summary_file.read_text()
# Check that page with metadata ignore is excluded from llms-full.txt
assert "Page Ignored by Metadata" not in llms_full_content
assert "This page should not appear in llms-full.txt" not in llms_full_content
# Check that page with metadata ignore is also excluded from llms.txt
# This should NOT contain a link to the ignored page
assert "Page Ignored by Metadata" not in llms_summary_content
assert "page_ignored_metadata.html" not in llms_summary_content
-146
View File
@@ -117,152 +117,6 @@ def test_max_lines_limit(temp_dir, rootdir):
app.docutils_conf_path.unlink() app.docutils_conf_path.unlink()
def test_on_exceed_skip(temp_dir, rootdir):
"""Test that skip action works when size limit is exceeded."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "skip-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "warn_skip",
},
)
app.build()
# Check that the output file was NOT created
output_file = Path(app.outdir) / "skip-test.txt"
assert (
not output_file.exists()
), f"Output file {output_file} should not exist with skip action"
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_on_exceed_keep(temp_dir, rootdir):
"""Test that keep action works when size limit is exceeded."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "keep-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "info_keep",
},
)
app.build()
# Check that the output file WAS created despite exceeding limit
output_file = Path(app.outdir) / "keep-test.txt"
assert (
output_file.exists()
), f"Output file {output_file} should exist with keep action"
# Verify it has content
content = output_file.read_text()
assert len(content) > 0, "Output file should have content with keep action"
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_on_exceed_note(temp_dir, rootdir):
"""Test that note action works when size limit is exceeded."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "note-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "warn_note",
},
)
app.build()
# Check that the output file WAS created with placeholder content
output_file = Path(app.outdir) / "note-test.txt"
assert (
output_file.exists()
), f"Output file {output_file} should exist with note action"
# Verify it has the placeholder content
content = output_file.read_text()
assert (
"This file was not generated because it exceeded the configured size limit."
in content
)
assert "llms_txt_full_max_size" in content
assert "llms_txt_full_size_policy" in content
assert "Configured max size: 20 lines" in content
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_on_exceed_invalid_config(temp_dir, rootdir):
"""Test behavior with invalid configuration values."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "invalid-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "invalid_format", # Invalid config
},
)
app.build()
# Should fall back to default behavior (warn_skip)
output_file = Path(app.outdir) / "invalid-test.txt"
assert (
not output_file.exists()
), f"Output file {output_file} should not exist with invalid config fallback"
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_title_override(temp_dir, rootdir): def test_title_override(temp_dir, rootdir):
"""Test that the title override works correctly.""" """Test that the title override works correctly."""
from sphinx.testing.util import SphinxTestApp from sphinx.testing.util import SphinxTestApp
-822
View File
@@ -41,42 +41,6 @@ def test_setup_returns_valid_dict():
assert "parallel_write_safe" in result assert "parallel_write_safe" in result
def test_builder_inited_with_disallowed_builder():
"""Test that disallowed builders do not trigger extension setup."""
import sphinx_llms_txt
# Reset global state
sphinx_llms_txt._manager = sphinx_llms_txt.LLMSFullManager()
sphinx_llms_txt._root_first_paragraph = ""
# Mock a Sphinx app with a disallowed builder
class MockBuilder:
name = "text" # Not in allowed list
class MockApp:
def __init__(self):
self.config_values = {}
self.connections = {}
self.builder = MockBuilder()
def add_config_value(self, name, default, rebuild):
self.config_values[name] = (default, rebuild)
def connect(self, event, handler):
self.connections[event] = handler
app = MockApp()
setup(app)
# Trigger builder-inited
builder_inited_handler = app.connections["builder-inited"]
builder_inited_handler(app)
# With disallowed builder, other events should NOT be connected
assert "doctree-resolved" not in app.connections
assert "build-finished" not in app.connections
def test_document_collector_initialization(): def test_document_collector_initialization():
"""Test initialization of DocumentCollector.""" """Test initialization of DocumentCollector."""
collector = DocumentCollector() collector = DocumentCollector()
@@ -370,789 +334,3 @@ def test_write_verbose_info_with_baseurl(tmp_path):
assert "- [Home Page](https://example.org/index.html)" in content assert "- [Home Page](https://example.org/index.html)" in content
assert "- [About Us](https://example.org/about.html)" in content assert "- [About Us](https://example.org/about.html)" in content
def test_get_source_suffixes_with_dict():
"""Test _get_source_suffixes method with dict source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with dict source_suffix
class MockApp:
class Config:
source_suffix = {".rst": None, ".md": None, ".txt": None}
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert set(suffixes) == {".rst", ".md", ".txt"}
def test_get_source_suffixes_with_list():
"""Test _get_source_suffixes method with list source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with list source_suffix
class MockApp:
class Config:
source_suffix = [".rst", ".md"]
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst", ".md"]
def test_get_source_suffixes_with_string():
"""Test _get_source_suffixes method with string source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with string source_suffix
class MockApp:
class Config:
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst"]
def test_get_source_suffixes_no_app():
"""Test _get_source_suffixes method with no app set."""
from sphinx_llms_txt.manager import LLMSFullManager
manager = LLMSFullManager()
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst"] # Default fallback
def test_html_sourcelink_suffix_default():
"""Test html_sourcelink_suffix defaults to .txt when no app is set."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
manager = LLMSFullManager()
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with default .txt suffix
test_file = f"{sources_dir}/index.rst.txt"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it uses .txt as the default suffix
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed (check if output file exists)
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_html_sourcelink_suffix_custom():
"""Test html_sourcelink_suffix uses custom value from Sphinx config."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with custom html_sourcelink_suffix
class MockApp:
class Config:
html_sourcelink_suffix = "source"
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with custom .source suffix
test_file = f"{sources_dir}/index.rst.source"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it uses .source as the custom suffix
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_html_sourcelink_suffix_with_dot():
"""Test html_sourcelink_suffix adds dot if missing."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with html_sourcelink_suffix without leading dot
class MockApp:
class Config:
html_sourcelink_suffix = "src" # No leading dot
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with .src suffix (dot should be added automatically)
test_file = f"{sources_dir}/index.rst.src"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it adds the dot and finds the file
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_mixed_source_file_formats():
"""Test handling of mixed source file formats (.rst, .md, .txt)."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with multiple source suffixes
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
source_suffix = {".rst": None, ".md": None, ".txt": None}
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create test source files with different formats
files_to_create = [
f"{sources_dir}/page1.rst.txt",
f"{sources_dir}/page2.md.txt",
f"{sources_dir}/page3.txt.txt",
]
for test_file in files_to_create:
with open(test_file, "w") as f:
f.write(f"Content for {os.path.basename(test_file)}")
# Mock env with all documents
class MockEnv:
all_docs = {"page1": None, "page2": None, "page3": None}
titles = {
"page1": type("TitleNode", (), {"astext": lambda: "Page 1"})(),
"page2": type("TitleNode", (), {"astext": lambda: "Page 2"})(),
"page3": type("TitleNode", (), {"astext": lambda: "Page 3"})(),
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("page1")
# Test that all file formats are found and processed
manager.combine_sources(outdir, srcdir)
# Verify the output file was created and contains content from all formats
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
with open(output_file, "r") as f:
content = f.read()
# Should contain content from all three files
assert "Content for page1.rst.txt" in content
assert "Content for page2.md.txt" in content
assert "Content for page3.txt.txt" in content
def test_source_suffix_detection_priority():
"""Test source suffix detection tries formats in correct order for docnames."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with ordered source suffixes
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
source_suffix = [".rst", ".md"] # rst has priority over md
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create both .rst and .md versions of the same document
# Only create files for the specific docname "index"
rst_file = f"{sources_dir}/index.rst.txt"
md_file = f"{sources_dir}/index.md.txt"
with open(rst_file, "w") as f:
f.write("RST content for index")
with open(md_file, "w") as f:
f.write("Markdown content for index")
# Mock env with only the index document
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Index Page"})()
}
toctree_includes = {"index": []}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test the priority behavior
manager.combine_sources(outdir, srcdir)
# Check that output file was created
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
with open(output_file, "r") as f:
content = f.read()
# The system should prefer RST over MD for the "index" docname
# But since both files exist and the second phase adds remaining files,
# both will be included. The test verifies that RST appears first
# (indicating it was found first in the priority order)
assert "RST content for index" in content
# Find positions to verify order
rst_pos = content.find("RST content for index")
md_pos = content.find("Markdown content for index")
# RST should come before MD (due to priority in toctree processing)
assert rst_pos < md_pos, "RST content should appear before MD content"
def test_summary_default_uses_first_paragraph():
"""
Test that summary defaults to first paragraph of root document when not configured.
"""
from docutils import nodes
from docutils.frontend import OptionParser
from docutils.parsers.rst import Parser
from docutils.utils import new_document
from sphinx_llms_txt import build_finished, doctree_resolved
# Create a proper document with settings
settings = OptionParser(components=(Parser,)).get_default_values()
doctree = new_document("<rst-doc>", settings)
title = nodes.title(text="Test Title")
paragraph = nodes.paragraph(
text="This is the first paragraph that should be used as summary."
)
doctree.append(title)
doctree.append(paragraph)
# Mock Sphinx app
class MockApp:
class Config:
master_doc = "index"
llms_txt_summary = None # Not configured
llms_txt_file = True
llms_txt_filename = "llms.txt"
llms_txt_title = None
llms_txt_full_file = True
llms_txt_full_filename = "llms-full.txt"
llms_txt_full_max_size = None
llms_txt_full_size_policy = "warn_skip"
llms_txt_directives = []
llms_txt_exclude = []
llms_txt_code_files = []
llms_txt_code_base_path = None
html_baseurl = ""
config = Config()
outdir = "/tmp/build"
srcdir = "/tmp/source"
class Env:
titles = {
"index": type("TitleNode", (), {"astext": lambda self: "Test Title"})()
}
env = Env()
app = MockApp()
# Reset the global state
import sphinx_llms_txt
sphinx_llms_txt._root_first_paragraph = ""
# Call doctree_resolved to extract the first paragraph
doctree_resolved(app, doctree, "index")
# Verify the first paragraph was extracted
assert (
sphinx_llms_txt._root_first_paragraph
== "This is the first paragraph that should be used as summary."
)
# Mock the manager methods to avoid actual file operations
original_combine_sources = sphinx_llms_txt._manager.combine_sources
sphinx_llms_txt._manager.combine_sources = lambda outdir, srcdir: None
# Call build_finished and verify the summary is set correctly
build_finished(app, None)
# Check that the summary was properly configured
assert (
sphinx_llms_txt._manager.config["llms_txt_summary"]
== "This is the first paragraph that should be used as summary."
)
# Restore original method
sphinx_llms_txt._manager.combine_sources = original_combine_sources
def test_code_files_include_exclude_patterns(tmp_path):
"""Test the +/- pattern syntax for llms_txt_code_files configuration."""
from sphinx_llms_txt.manager import LLMSFullManager
# Create test directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
docs_dir = src_dir / "docs"
docs_dir.mkdir()
cache_dir = docs_dir / "__pycache__"
cache_dir.mkdir()
# Create test files
(docs_dir / "example.rst").write_text("Example RST content")
(docs_dir / "guide.rst").write_text("Guide RST content")
(docs_dir / "backup.bak").write_text("Backup file content")
(cache_dir / "compiled.pyc").write_text("Compiled Python")
# Create manager and set source directory
manager = LLMSFullManager()
manager.srcdir = str(src_dir)
# Test configuration with include/exclude patterns
config = {
"llms_txt_code_files": [
"+:docs/**/*.rst", # Include all RST files in docs
"-:docs/**/__pycache__/**", # Exclude pycache files
"-:docs/**/*.bak", # Exclude backup files
]
}
manager.set_config(config)
# Process code files
code_parts, _ = manager._process_code_files()
# Verify we have the expected number of files
assert len(code_parts) == 2, f"Expected 2 files, got {len(code_parts)}"
# Extract file titles from code blocks
titles = []
for part in code_parts:
lines = part.strip().split("\n")
if lines:
titles.append(lines[0])
# Verify expected files are included
assert "docs/example.rst" in titles
assert "docs/guide.rst" in titles
# Verify excluded files are not present
content = "\n".join(code_parts)
assert "backup.bak" not in content
assert "__pycache__" not in content
assert "compiled.pyc" not in content
def test_code_files_exclude_only_patterns(tmp_path):
"""Test that exclude-only patterns result in no files being included."""
from sphinx_llms_txt.manager import LLMSFullManager
# Create test directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
docs_dir = src_dir / "docs"
docs_dir.mkdir()
# Create test files
(docs_dir / "example.rst").write_text("Example RST content")
# Create manager and set source directory
manager = LLMSFullManager()
manager.srcdir = str(src_dir)
# Test configuration with only exclude patterns
config = {
"llms_txt_code_files": [
"-:docs/**/*.rst", # Only exclude pattern, no includes
]
}
manager.set_config(config)
# Process code files
code_parts, _ = manager._process_code_files()
# Should have no files with exclude-only patterns
assert len(code_parts) == 0, "Should have no files with exclude-only patterns"
def test_code_files_no_prefix_patterns(tmp_path):
"""Test that patterns without prefix are ignored."""
from sphinx_llms_txt.manager import LLMSFullManager
# Create test directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
docs_dir = src_dir / "docs"
docs_dir.mkdir()
# Create test files
(docs_dir / "example.rst").write_text("Example RST content")
(docs_dir / "backup.bak").write_text("Backup file content")
# Create manager and set source directory
manager = LLMSFullManager()
manager.srcdir = str(src_dir)
# Test configuration with no prefix (should be ignored)
config = {
"llms_txt_code_files": [
"docs/**/*.rst", # No prefix = ignored
"+:docs/**/*.rst", # Include RST files
"-:docs/**/*.bak", # Exclude backup files
]
}
manager.set_config(config)
# Process code files
code_parts, _ = manager._process_code_files()
# Should include RST files (from +: pattern) and exclude BAK files (from -: pattern)
assert len(code_parts) == 1, "Should include RST files and exclude BAK files"
content = "\n".join(code_parts)
assert "Example RST content" in content
assert "backup.bak" not in content
def test_code_files_ignored_patterns(tmp_path, caplog):
"""Test that patterns without +: or -: prefix log a warning and are ignored."""
from unittest.mock import patch
from sphinx_llms_txt.manager import LLMSFullManager
# Create test directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
docs_dir = src_dir / "docs"
docs_dir.mkdir()
# Create test files
(docs_dir / "example.rst").write_text("Example RST content")
# Use a mock to capture the warning message directly
captured_warnings = []
def capture_warning(message, *args, **kwargs):
captured_warnings.append(message)
# Patch the logger to capture warnings
with patch("sphinx_llms_txt.manager.logger.warning", side_effect=capture_warning):
# Create manager and set source directory
manager = LLMSFullManager()
manager.srcdir = str(src_dir)
# Test configuration with only no-prefix patterns (should result in no files)
config = {
"llms_txt_code_files": [
"docs/**/*.rst", # No prefix = ignored with warning
]
}
manager.set_config(config)
# Process code files
code_parts, _ = manager._process_code_files()
# Should have no files since the pattern without prefix is ignored
assert (
len(code_parts) == 0
), "Should have no files when only using patterns without prefix"
# Check that a warning was logged
assert (
len(captured_warnings) == 1
), f"Expected 1 warning, got {len(captured_warnings)}"
assert (
"Code file pattern 'docs/**/*.rst' ignored." in captured_warnings[0]
), f"Warning message should contain expected text. Got: {captured_warnings[0]}"
def test_llms_txt_generated_without_sources_dir(tmp_path):
"""Test that llms.txt is generated even when _sources directory doesn't exist."""
from sphinx_llms_txt.manager import LLMSFullManager
# Create manager
manager = LLMSFullManager()
# Set config to enable llms.txt
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_full_file": True,
"llms_txt_full_filename": "llms-full.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
manager.set_config(config)
# Create directories (but no _sources)
outdir = tmp_path / "build"
srcdir = tmp_path / "source"
outdir.mkdir()
srcdir.mkdir()
# Mock env with documents
class MockEnv:
all_docs = {"index": None, "about": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda self: "Home"})(),
"about": type("TitleNode", (), {"astext": lambda self: "About"})(),
}
toctree_includes = {"index": ["about"]}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Update page titles directly in the collector
manager.update_page_title("index", "Home")
manager.update_page_title("about", "About")
# Call combine_sources - should generate llms.txt even without _sources
manager.combine_sources(str(outdir), str(srcdir))
# Verify llms.txt was created
llms_txt = outdir / "llms.txt"
assert llms_txt.exists(), "llms.txt should be generated even without _sources"
# Verify llms-full.txt was NOT created (since no _sources)
llms_full_txt = outdir / "llms-full.txt"
assert (
not llms_full_txt.exists()
), "llms-full.txt should not be generated without _sources"
# Read llms.txt and verify it has content
with open(llms_txt, "r", encoding="utf-8") as f:
content = f.read()
# Should contain page titles and links
assert "Home" in content
assert "About" in content
assert "index.html" in content
assert "about.html" in content
def test_llms_txt_no_warning_when_full_file_disabled(tmp_path, caplog):
"""
Test that no warning is logged when llms_txt_full_file=False and
_sources doesn't exist.
"""
from unittest.mock import patch
from sphinx_llms_txt.manager import LLMSFullManager
# Create manager
manager = LLMSFullManager()
# Set config with llms_txt_full_file=False
config = {
"llms_txt_file": True,
"llms_txt_filename": "llms.txt",
"llms_txt_full_file": False, # User doesn't want llms-full.txt
"llms_txt_full_filename": "llms-full.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
manager.set_config(config)
# Create directories (but no _sources)
outdir = tmp_path / "build"
srcdir = tmp_path / "source"
outdir.mkdir()
srcdir.mkdir()
# Mock env with documents
class MockEnv:
all_docs = {"index": None}
titles = {"index": type("TitleNode", (), {"astext": lambda self: "Home"})()}
toctree_includes = {"index": []}
manager.set_env(MockEnv())
manager.set_master_doc("index")
manager.update_page_title("index", "Home")
# Capture warnings
captured_warnings = []
def capture_warning(message, *args, **kwargs):
if "_sources" in str(message):
captured_warnings.append(message)
with patch("sphinx_llms_txt.manager.logger.warning", side_effect=capture_warning):
# Call combine_sources
manager.combine_sources(str(outdir), str(srcdir))
# Verify NO warning was logged since llms_txt_full_file=False
assert (
len(captured_warnings) == 0
), "No warning should be logged when llms_txt_full_file=False"
# Verify llms.txt was still created
llms_txt = outdir / "llms.txt"
assert llms_txt.exists()
+3 -156
View File
@@ -102,7 +102,7 @@ def test_process_path_directives_with_html_baseurl(tmp_path):
def test_process_path_directives_absolute_urls(tmp_path): def test_process_path_directives_absolute_urls(tmp_path):
"""Test that absolute URLs are not modified but absolute paths get base URL.""" """Test that absolute URLs are not modified."""
# Create a processor # Create a processor
config = { config = {
"llms_txt_directives": [], "llms_txt_directives": [],
@@ -127,17 +127,10 @@ def test_process_path_directives_absolute_urls(tmp_path):
with open(source_file, "w", encoding="utf-8") as f: with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content) f.write(source_content)
# Process the directives # Process the directives (should remain unchanged)
processed_content = processor._process_path_directives(source_content, source_file) processed_content = processor._process_path_directives(source_content, source_file)
# Expected: URLs and data URIs unchanged, absolute paths get base URL assert processed_content == source_content
expected_content = (
".. image:: https://othersite.com/images/test.png\n"
".. image:: https://example.com/docs/absolute/path/image.png\n"
".. image:: data:image/png;base64,iVBORw0KG...\n"
)
assert processed_content == expected_content
def test_process_path_directives_custom_directives(tmp_path): def test_process_path_directives_custom_directives(tmp_path):
@@ -258,149 +251,3 @@ def test_process_content_end_to_end(tmp_path):
) )
assert processed_content == expected_content assert processed_content == expected_content
def test_process_path_directives_images_directory(tmp_path):
"""Test that _images directory paths are handled correctly."""
# Create a processor with base URL
config = {
"llms_txt_directives": [],
"html_baseurl": "https://example.com/docs",
}
processor = DocumentProcessor(config)
# Create source directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
processor.srcdir = str(src_dir)
# Create _sources directory to mimic Sphinx output
build_dir = tmp_path / "build"
build_dir.mkdir()
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Create a source file with various _images directory paths
source_content = (
"Some content.\n"
".. image:: _images/test.png\n" # Relative _images should become /_images
".. image:: /_images/absolute.png\n" # Absolute _images should get base URL
".. figure:: _images/figure.png\n" # Test with figure directive too
" :alt: A test figure\n"
".. image:: images/normal.png\n" # Normal relative path should be unchanged
)
# Create source file in sources directory to simulate Sphinx build output
source_file = sources_dir / "page.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
# Expected: _images paths should be converted and get base URL
expected_content = (
"Some content.\n"
".. image:: https://example.com/docs/_images/test.png\n"
".. image:: https://example.com/docs/_images/absolute.png\n"
".. figure:: https://example.com/docs/_images/figure.png\n"
" :alt: A test figure\n"
".. image:: https://example.com/docs/images/normal.png\n"
)
assert processed_content == expected_content
def test_process_path_directives_images_directory_no_baseurl(tmp_path):
"""
Test that _images directory paths work correctly without base URL.
Only converts when image exists.
"""
# Create a processor without base URL
config = {
"llms_txt_directives": [],
"html_baseurl": "",
}
processor = DocumentProcessor(config)
# Create source directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
processor.srcdir = str(src_dir)
# Create _sources directory to mimic Sphinx output
build_dir = tmp_path / "build"
build_dir.mkdir()
sources_dir = build_dir / "_sources"
sources_dir.mkdir()
# Create _images directory and one test image
images_dir = build_dir / "_images"
images_dir.mkdir()
(images_dir / "test.png").write_text("fake image content")
# Note: absolute.png is not created, so it won't be converted
# Create a source file with _images directory paths
source_content = (
".. image:: _images/test.png\n" # Should become /_images (image exists)
".. image:: /_images/absolute.png\n" # Should stay unchanged (absolute path)
)
# Create source file in sources directory to simulate Sphinx build output
source_file = sources_dir / "page.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
# Expected: only test.png gets converted because it exists in _images
expected_content = (
".. image:: /_images/test.png\n" # Converted because image exists
".. image:: /_images/absolute.png\n" # Absolute path unchanged
)
assert processed_content == expected_content
def test_process_path_directives_all_absolute_paths_get_baseurl(tmp_path):
"""Test that all absolute paths (starting with /) get base URL prepended."""
# Create a processor with base URL
config = {
"llms_txt_directives": [],
"html_baseurl": "https://mysite.com/docs/",
}
processor = DocumentProcessor(config)
# Create source directory structure
src_dir = tmp_path / "src"
src_dir.mkdir()
processor.srcdir = str(src_dir)
# Create a source file with various absolute paths
source_content = (
".. image:: /static/images/logo.png\n"
".. figure:: /assets/diagrams/flow.svg\n"
".. image:: /media/photos/team.jpg\n"
" :alt: Team photo\n"
".. image:: relative/path.png\n" # This should still get normal processing
)
# Create source file
source_file = src_dir / "page.txt"
with open(source_file, "w", encoding="utf-8") as f:
f.write(source_content)
# Process the directives
processed_content = processor._process_path_directives(source_content, source_file)
# Expected: All absolute paths get base URL prepended
expected_content = (
".. image:: https://mysite.com/docs/static/images/logo.png\n"
".. figure:: https://mysite.com/docs/assets/diagrams/flow.svg\n"
".. image:: https://mysite.com/docs/media/photos/team.jpg\n"
" :alt: Team photo\n"
".. image:: https://mysite.com/docs/relative/path.png\n"
)
assert processed_content == expected_content