Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8625868df7 | ||
|
|
b2aa7897d3 | ||
|
|
12d695015e | ||
|
|
29c400e122 | ||
|
|
834a57a158 | ||
|
|
efbe8e0cda | ||
|
|
68860c7dda | ||
|
|
055261b3bb | ||
|
|
a274621b32 | ||
|
|
2f62461703 | ||
|
|
da48421b03 | ||
|
|
9385670cfe |
@@ -1,72 +1,6 @@
|
|||||||
Changelog
|
Changelog
|
||||||
=========
|
=========
|
||||||
|
|
||||||
0.5.2
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Remove support for singlehtml
|
|
||||||
`#40 <https://github.com/jdillard/sphinx-llms-txt/pull/40>`_
|
|
||||||
|
|
||||||
0.5.1
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Only allow builders that have _sources directory
|
|
||||||
`#38 <https://github.com/jdillard/sphinx-llms-txt/pull/38>`_
|
|
||||||
|
|
||||||
0.5.0
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Add :ref:`block_level_ignore` and :ref:`page_level_ignore`
|
|
||||||
`#33 <https://github.com/jdillard/sphinx-llms-txt/pull/33>`_
|
|
||||||
- Add :confval:`llms_txt_full_size_policy` configuration option to control behavior when :confval:`llms_txt_full_max_size` is exceeded.
|
|
||||||
`#35 <https://github.com/jdillard/sphinx-llms-txt/pull/35>`_
|
|
||||||
|
|
||||||
0.4.1
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Fix include paths and spacing
|
|
||||||
`#31 <https://github.com/jdillard/sphinx-llms-txt/pull/31>`_
|
|
||||||
|
|
||||||
0.4.0
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Add support for including source code files with :confval:`llms_txt_code_files` and :confval:`llms_txt_code_base_path` configuration options
|
|
||||||
`#24 <https://github.com/jdillard/sphinx-llms-txt/pull/24>`_
|
|
||||||
|
|
||||||
0.3.2
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Fix image paths to deployed images
|
|
||||||
`#30 <https://github.com/jdillard/sphinx-llms-txt/pull/30>`_
|
|
||||||
|
|
||||||
0.3.1
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Fix issue when ``source_suffix`` equals ``source_link_suffix``
|
|
||||||
`#29 <https://github.com/jdillard/sphinx-llms-txt/pull/29>`_
|
|
||||||
|
|
||||||
0.3.0
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Use first paragraph as default for ``llms_txt_summary``
|
|
||||||
`#22 <https://github.com/jdillard/sphinx-llms-txt/pull/22>`_
|
|
||||||
|
|
||||||
0.2.4
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Support source file suffix detection
|
|
||||||
`#21 <https://github.com/jdillard/sphinx-llms-txt/pull/21>`_
|
|
||||||
|
|
||||||
0.2.3
|
|
||||||
-----
|
|
||||||
|
|
||||||
- Remove ``get_and_resolve_toctree`` method
|
|
||||||
`#19 <https://github.com/jdillard/sphinx-llms-txt/pull/19>`_
|
|
||||||
- Simplify ``_sources`` lookup
|
|
||||||
`#18 <https://github.com/jdillard/sphinx-llms-txt/pull/18>`_
|
|
||||||
- Add sphinx docs
|
|
||||||
`#16 <https://github.com/jdillard/sphinx-llms-txt/pull/16>`_
|
|
||||||
|
|
||||||
0.2.2
|
0.2.2
|
||||||
-----
|
-----
|
||||||
|
|
||||||
|
|||||||
@@ -1,20 +1,14 @@
|
|||||||
# Sphinx llms.txt generator
|
# Sphinx llms.txt generator
|
||||||
|
|
||||||
A Sphinx extension that generates a summary `llms.txt` file and a single combined documentation `llms-full.txt` file.
|
A Sphinx extension that generates a summary `llms.txt` file, written in Markdown, and a single combined documentation `llms-full.txt` file, written in reStructuredText.
|
||||||
|
|
||||||
[](https://pypi.python.org/pypi/sphinx-llms-txt)
|
[](https://pypi.python.org/pypi/sphinx-llms-txt)
|
||||||
[](https://anaconda.org/conda-forge/sphinx-llms-txt)
|
|
||||||
[](https://pepy.tech/project/sphinx-llms-txt)
|
[](https://pepy.tech/project/sphinx-llms-txt)
|
||||||
[](#)
|
|
||||||
|
|
||||||
## Documentation
|
## Documentation
|
||||||
|
|
||||||
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
|
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
|
||||||
|
|
||||||
## Contributing
|
|
||||||
|
|
||||||
Pull Requests welcome! See [Contributing](https://sphinx-llms-txt.readthedocs.io/en/latest/contributing.html) for instructions on how best to contribute.
|
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
MIT License - see LICENSE file for details.
|
MIT License - see LICENSE file for details.
|
||||||
|
|||||||
@@ -76,26 +76,15 @@ Handling Large Documentation
|
|||||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
|
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
|
||||||
You can set a maximum line count and control what happens when that limit is exceeded:
|
You can set a maximum line count:
|
||||||
|
|
||||||
.. code-block:: python
|
.. code-block:: python
|
||||||
|
|
||||||
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
|
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
|
||||||
llms_txt_full_size_policy = "warn_skip" # Default behavior
|
|
||||||
|
|
||||||
The ``llms_txt_full_size_policy`` setting controls both the log level and action taken when the size limit is exceeded.
|
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
|
||||||
It uses the format ``"<loglevel>_<action>"``:
|
|
||||||
|
|
||||||
**Log levels:**
|
.. tip:: Use :ref:`excluding_content` to remove less relevant pages.
|
||||||
- ``warn``: Log as a warning (default)
|
|
||||||
- ``info``: Log as informational message
|
|
||||||
|
|
||||||
**Actions:**
|
|
||||||
- ``skip``: Don't create the file (default)
|
|
||||||
- ``keep``: Create the file anyway, ignoring the size limit
|
|
||||||
- ``note``: Create a placeholder file explaining why the full file wasn't generated
|
|
||||||
|
|
||||||
.. tip:: Use :ref:`excluding_content` to remove less relevant pages and reduce the file size.
|
|
||||||
|
|
||||||
.. _custom_directive_handling:
|
.. _custom_directive_handling:
|
||||||
|
|
||||||
@@ -124,13 +113,6 @@ This ensures that paths in your custom directives are properly resolved in the g
|
|||||||
Excluding Content
|
Excluding Content
|
||||||
^^^^^^^^^^^^^^^^^
|
^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
There are several ways to exclude content from the generated ``llms-full.txt`` file:
|
|
||||||
|
|
||||||
.. _global_exclusion:
|
|
||||||
|
|
||||||
Global Page Exclusion
|
|
||||||
~~~~~~~~~~~~~~~~~~~~~~
|
|
||||||
|
|
||||||
You can exclude specific pages from being included in the generated files:
|
You can exclude specific pages from being included in the generated files:
|
||||||
|
|
||||||
.. code-block:: python
|
.. code-block:: python
|
||||||
@@ -142,111 +124,6 @@ You can exclude specific pages from being included in the generated files:
|
|||||||
]
|
]
|
||||||
|
|
||||||
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
|
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
|
||||||
It can also be used to reduce the size of llms-full.txt.
|
|
||||||
|
|
||||||
.. _page_level_ignore:
|
|
||||||
|
|
||||||
Page-Level Ignore Metadata
|
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
|
||||||
|
|
||||||
You can exclude individual pages by adding metadata at the top of any reStructuredText file:
|
|
||||||
|
|
||||||
.. code-block:: restructuredtext
|
|
||||||
|
|
||||||
:llms-txt-ignore: true
|
|
||||||
|
|
||||||
Page Title
|
|
||||||
==========
|
|
||||||
|
|
||||||
This entire page will be excluded from llms-full.txt
|
|
||||||
|
|
||||||
When this metadata is present, the entire page is skipped during processing.
|
|
||||||
|
|
||||||
.. _block_level_ignore:
|
|
||||||
|
|
||||||
Block-Level Ignore Directives
|
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
|
||||||
|
|
||||||
You can exclude specific sections within a page using ignore directives:
|
|
||||||
|
|
||||||
.. code-block:: restructuredtext
|
|
||||||
|
|
||||||
Page Title
|
|
||||||
==========
|
|
||||||
|
|
||||||
This content will be included in llms-full.txt.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
This content will be excluded from llms-full.txt.
|
|
||||||
|
|
||||||
Section To Ignore
|
|
||||||
-----------------
|
|
||||||
|
|
||||||
This entire section and any nested content will be ignored.
|
|
||||||
|
|
||||||
.. code-block:: python
|
|
||||||
|
|
||||||
# This code block will also be ignored
|
|
||||||
def ignored_function():
|
|
||||||
pass
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
This content will be included again.
|
|
||||||
|
|
||||||
Block-level ignores can be useful for:
|
|
||||||
|
|
||||||
- Removing internal notes or TODOs
|
|
||||||
- Hiding implementation details while keeping user-facing documentation
|
|
||||||
|
|
||||||
.. note::
|
|
||||||
- Multiple ignore blocks can be used within the same file
|
|
||||||
- Ignore directives work with any indentation level
|
|
||||||
|
|
||||||
.. _including_code_files:
|
|
||||||
|
|
||||||
Including Source Code Files
|
|
||||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
|
||||||
|
|
||||||
You can include source code files from your project at the end of :confval:`llms_txt_full_filename`.
|
|
||||||
|
|
||||||
Use include/exclude syntax to precisely control which files are included:
|
|
||||||
|
|
||||||
.. code-block:: python
|
|
||||||
|
|
||||||
llms_txt_code_files = [
|
|
||||||
"+:src/**/*.py", # Include all Python files in src
|
|
||||||
"-:src/**/__pycache__/**", # Exclude Python cache files
|
|
||||||
]
|
|
||||||
|
|
||||||
Pattern syntax:
|
|
||||||
|
|
||||||
- **+:pattern**: Include files matching the pattern. Processed first to collect matching files.
|
|
||||||
- **-:pattern**: Exclude files matching the pattern. Applied to filter out unwanted files.
|
|
||||||
|
|
||||||
Code files are processed as follows:
|
|
||||||
|
|
||||||
- **Glob patterns**: Use standard glob patterns (``*``, ``**``, ``?``) to match files
|
|
||||||
- **Relative paths**: Patterns are resolved relative to your Sphinx source directory
|
|
||||||
- **Formatting**: Each file is presented with a title and syntax-highlighted code block
|
|
||||||
|
|
||||||
.. _customizing_code_paths:
|
|
||||||
|
|
||||||
Customizing Code File Paths
|
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
|
||||||
|
|
||||||
By default, the extension automatically detects the relative path from your Sphinx source directory to the git root and strips that prefix from displayed file paths. You can customize this behavior:
|
|
||||||
|
|
||||||
.. code-block:: python
|
|
||||||
|
|
||||||
# Manually specify base path to strip
|
|
||||||
llms_txt_code_base_path = "../../"
|
|
||||||
|
|
||||||
# Disable path stripping entirely
|
|
||||||
llms_txt_code_base_path = ""
|
|
||||||
|
|
||||||
This helps create cleaner, more readable file paths in the generated documentation.
|
|
||||||
|
|
||||||
.. _using_html_baseurl:
|
.. _using_html_baseurl:
|
||||||
|
|
||||||
@@ -269,7 +146,7 @@ Integration Examples
|
|||||||
Complete Configuration Example
|
Complete Configuration Example
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
Here's a complete example showing multiple :doc:`configuration-values`:
|
Here's a complete example showing multiple configuration options:
|
||||||
|
|
||||||
.. code-block:: python
|
.. code-block:: python
|
||||||
|
|
||||||
@@ -277,7 +154,6 @@ Here's a complete example showing multiple :doc:`configuration-values`:
|
|||||||
llms_txt_filename = "ai-summary.txt"
|
llms_txt_filename = "ai-summary.txt"
|
||||||
llms_txt_full_filename = "ai-full-docs.txt"
|
llms_txt_full_filename = "ai-full-docs.txt"
|
||||||
llms_txt_full_max_size = 50000
|
llms_txt_full_max_size = 50000
|
||||||
llms_txt_full_size_policy = "warn_note"
|
|
||||||
|
|
||||||
# Content customization
|
# Content customization
|
||||||
llms_txt_title = "Project Documentation for AI Assistants"
|
llms_txt_title = "Project Documentation for AI Assistants"
|
||||||
@@ -292,11 +168,3 @@ Here's a complete example showing multiple :doc:`configuration-values`:
|
|||||||
|
|
||||||
# Content filtering
|
# Content filtering
|
||||||
llms_txt_exclude = ["search", "genindex", "404", "private_*"]
|
llms_txt_exclude = ["search", "genindex", "404", "private_*"]
|
||||||
|
|
||||||
# Source code inclusion with include/exclude patterns
|
|
||||||
llms_txt_code_files = [
|
|
||||||
"+:../../src/**/*.py", # Include Python files
|
|
||||||
"+:../../config/*.yaml", # Include config files
|
|
||||||
"-:../../src/**/__pycache__/**", # Exclude cache files
|
|
||||||
]
|
|
||||||
llms_txt_code_base_path = "../../"
|
|
||||||
|
|||||||
+1
-6
@@ -15,7 +15,6 @@ import subprocess
|
|||||||
project = "sphinx-llms-txt"
|
project = "sphinx-llms-txt"
|
||||||
copyright = "Jared Dillard"
|
copyright = "Jared Dillard"
|
||||||
author = "Jared Dillard"
|
author = "Jared Dillard"
|
||||||
llms_txt_code_files = ["+:../../sphinx_llms_txt/*.py"]
|
|
||||||
llms_txt_summary = """
|
llms_txt_summary = """
|
||||||
A Sphinx extension that generates a summary llms.txt file,written in Markdown,
|
A Sphinx extension that generates a summary llms.txt file,written in Markdown,
|
||||||
and a single combined documentation llms-full.txt file, written in reStructuredText.
|
and a single combined documentation llms-full.txt file, written in reStructuredText.
|
||||||
@@ -82,11 +81,7 @@ html_theme = "furo"
|
|||||||
# further. For a list of options available for each theme, see the
|
# further. For a list of options available for each theme, see the
|
||||||
# documentation.
|
# documentation.
|
||||||
#
|
#
|
||||||
html_theme_options = {
|
html_theme_options = {}
|
||||||
"source_repository": "https://github.com/jdillard/sphinx-llms-txt/",
|
|
||||||
"source_branch": "main",
|
|
||||||
"source_directory": "docs/source/",
|
|
||||||
}
|
|
||||||
|
|
||||||
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/"
|
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/"
|
||||||
|
|
||||||
|
|||||||
@@ -24,22 +24,11 @@ Project Configuration Values
|
|||||||
- **Type**: integer or ``None``
|
- **Type**: integer or ``None``
|
||||||
- **Default**: ``None`` (no limit)
|
- **Default**: ``None`` (no limit)
|
||||||
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
||||||
Behavior when exceeded is controlled by :confval:`llms_txt_full_size_policy`.
|
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
||||||
See :ref:`handling_large_documentation`.
|
See :ref:`handling_large_documentation`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
.. confval:: llms_txt_full_size_policy
|
|
||||||
|
|
||||||
- **Type**: string
|
|
||||||
- **Default**: ``'warn_skip'``
|
|
||||||
- **Description**: Controls what happens when :confval:`llms_txt_full_max_size` is exceeded.
|
|
||||||
Format is ``<loglevel>_<action>``. Log levels: ``warn``, ``info``.
|
|
||||||
Actions: ``skip``, ``keep``, ``note``.
|
|
||||||
See :ref:`handling_large_documentation`.
|
|
||||||
|
|
||||||
.. versionadded:: 0.5.0
|
|
||||||
|
|
||||||
.. confval:: llms_txt_file
|
.. confval:: llms_txt_file
|
||||||
|
|
||||||
- **Type**: boolean
|
- **Type**: boolean
|
||||||
@@ -78,8 +67,8 @@ Project Configuration Values
|
|||||||
|
|
||||||
.. confval:: llms_txt_summary
|
.. confval:: llms_txt_summary
|
||||||
|
|
||||||
- **Type**: string
|
- **Type**: string or ``None``
|
||||||
- **Default**: The first paragraph in the root document, else an empty string
|
- **Default**: ``None``
|
||||||
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
|
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
|
||||||
See :ref:`custom_summary`.
|
See :ref:`custom_summary`.
|
||||||
|
|
||||||
@@ -89,26 +78,7 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: list of strings
|
- **Type**: list of strings
|
||||||
- **Default**: ``[]``
|
- **Default**: ``[]``
|
||||||
- **Description**: A list of pages to ignore using glob patterns.
|
- **Description**: A list of pages to ignore.
|
||||||
See :ref:`excluding_content`.
|
See :ref:`excluding_content`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.1
|
.. versionadded:: 0.2.1
|
||||||
|
|
||||||
.. confval:: llms_txt_code_files
|
|
||||||
|
|
||||||
- **Type**: list of strings
|
|
||||||
- **Default**: ``[]``
|
|
||||||
- **Description**: A list of glob patterns that appends source code files to :confval:`llms_txt_full_filename`.
|
|
||||||
See :ref:`including_code_files`.
|
|
||||||
|
|
||||||
.. versionadded:: 0.4.0
|
|
||||||
|
|
||||||
.. confval:: llms_txt_code_base_path
|
|
||||||
|
|
||||||
- **Type**: string or ``None``
|
|
||||||
- **Default**: ``None`` (auto-detect from git root)
|
|
||||||
- **Description**: Base path to strip from code file paths when displaying titles.
|
|
||||||
When ``None``, automatically detects the relative path from the Sphinx source
|
|
||||||
directory to the git root and strips that prefix from file paths.
|
|
||||||
|
|
||||||
.. versionadded:: 0.4.0
|
|
||||||
|
|||||||
@@ -1,6 +1,11 @@
|
|||||||
Getting Started
|
Getting Started
|
||||||
===============
|
===============
|
||||||
|
|
||||||
|
Demo
|
||||||
|
----
|
||||||
|
|
||||||
|
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
||||||
|
|
||||||
Installation
|
Installation
|
||||||
------------
|
------------
|
||||||
|
|
||||||
@@ -10,12 +15,6 @@ Directly install via ``pip`` by using:
|
|||||||
|
|
||||||
pip install sphinx-llms-txt
|
pip install sphinx-llms-txt
|
||||||
|
|
||||||
Or with ``conda`` via ``conda-forge``:
|
|
||||||
|
|
||||||
.. code::
|
|
||||||
|
|
||||||
conda install -c conda-forge sphinx-llms-txt
|
|
||||||
|
|
||||||
Usage
|
Usage
|
||||||
-----
|
-----
|
||||||
|
|
||||||
@@ -27,12 +26,25 @@ Add the extension to your Sphinx configuration (``conf.py``):
|
|||||||
'sphinx_llms_txt',
|
'sphinx_llms_txt',
|
||||||
]
|
]
|
||||||
|
|
||||||
After the HTML finishes building, **sphinx-llms-txt** will output the location of the output files::
|
Once added, the extension will automatically generate the LLMs.txt files during the build process.
|
||||||
|
|
||||||
sphinx-llms-txt: Created /path/to/_build/html/llms-full.txt with 45 sources and 6879 lines
|
|
||||||
sphinx-llms-txt: created /path/to/_build/html/llms.txt
|
|
||||||
|
|
||||||
|
|
||||||
.. tip:: Make sure to confirm the accuracy of the output files after installs and upgrades.
|
|
||||||
|
|
||||||
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
|
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
|
||||||
|
|
||||||
|
How It Works
|
||||||
|
-----------
|
||||||
|
|
||||||
|
During the Sphinx build process:
|
||||||
|
|
||||||
|
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
|
||||||
|
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
|
||||||
|
3. **Path Resolution**: Transforms relative paths in directives to full paths
|
||||||
|
4. **Output Generation**: Creates two optional files:
|
||||||
|
|
||||||
|
- ``llms.txt``: A concise summary of your documentation, in Markdown
|
||||||
|
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
|
||||||
|
|
||||||
|
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
|
||||||
|
|
||||||
|
|
||||||
|
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
||||||
|
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
||||||
|
|||||||
+1
-34
@@ -3,26 +3,7 @@ Sphinx llms.txt Generator
|
|||||||
|
|
||||||
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|
||||||
|
|
||||||
|PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars|
|
|PyPI version|
|
||||||
|
|
||||||
Demo
|
|
||||||
----
|
|
||||||
|
|
||||||
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
|
||||||
|
|
||||||
Highlights
|
|
||||||
----------
|
|
||||||
|
|
||||||
1. **Content Collection**: Quickly gathers content from _sources, without needing a separate build
|
|
||||||
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
|
|
||||||
3. **Path Resolution**: Transforms relative paths in directives to full paths
|
|
||||||
4. **Output Generation**: Creates two optional files:
|
|
||||||
|
|
||||||
- ``llms.txt``: A concise summary of your documentation, in Markdown
|
|
||||||
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
|
|
||||||
|
|
||||||
5. **Content Filtering**: Allows you to exclude specific pages or sections
|
|
||||||
6. **Source Code**: Allows you to include specific source code files
|
|
||||||
|
|
||||||
.. toctree::
|
.. toctree::
|
||||||
:maxdepth: 2
|
:maxdepth: 2
|
||||||
@@ -34,22 +15,8 @@ Highlights
|
|||||||
changelog
|
changelog
|
||||||
|
|
||||||
|
|
||||||
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
|
||||||
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
|
||||||
.. _Sphinx: http://sphinx-doc.org/
|
.. _Sphinx: http://sphinx-doc.org/
|
||||||
|
|
||||||
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
||||||
:target: https://pypi.python.org/pypi/sphinx-llms-txt
|
:target: https://pypi.python.org/pypi/sphinx-llms-txt
|
||||||
:alt: Latest PyPi Version
|
:alt: Latest PyPi Version
|
||||||
.. |Conda Version| image:: https://img.shields.io/conda/vn/conda-forge/sphinx-llms-txt.svg
|
|
||||||
:target: https://anaconda.org/conda-forge/sphinx-llms-txt
|
|
||||||
:alt: Latest Conda Version
|
|
||||||
.. |Downloads| image:: https://static.pepy.tech/badge/sphinx-llms-txt/month
|
|
||||||
:target: https://pepy.tech/project/sphinx-llms-txt
|
|
||||||
:alt: PyPi Downloads per month
|
|
||||||
.. |Parallel Safe| image:: https://img.shields.io/badge/parallel%20safe-true-brightgreen
|
|
||||||
:target: #
|
|
||||||
:alt: Parallel read/write safe
|
|
||||||
.. |GitHub Stars| image:: https://img.shields.io/github/stars/jdillard/sphinx-llms-txt?style=social
|
|
||||||
:target: https://github.com/jdillard/sphinx-llms-txt
|
|
||||||
:alt: GitHub Repository stars
|
|
||||||
|
|||||||
+10
-56
@@ -1,14 +1,5 @@
|
|||||||
"""
|
"""
|
||||||
Sphinx extension that generates llms.txt and llms-full.txt files for LLM consumption.
|
Sphinx extension to create a combined sources file (llms-full.txt)
|
||||||
|
|
||||||
This extension collects documentation content from Sphinx projects and generates
|
|
||||||
two output files:
|
|
||||||
- llms.txt: A concise Markdown summary with project overview and page links
|
|
||||||
- llms-full.txt: A comprehensive reStructuredText file containing all documentation
|
|
||||||
content with resolved includes and path references
|
|
||||||
|
|
||||||
The extension processes content during the build phase, handles page-level and
|
|
||||||
block-level ignore directives, and can optionally include source code files.
|
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from typing import Any, Dict
|
from typing import Any, Dict
|
||||||
@@ -21,7 +12,7 @@ from .manager import LLMSFullManager
|
|||||||
from .processor import DocumentProcessor
|
from .processor import DocumentProcessor
|
||||||
from .writer import FileWriter
|
from .writer import FileWriter
|
||||||
|
|
||||||
__version__ = "0.5.2"
|
__version__ = "0.2.2"
|
||||||
|
|
||||||
# Export classes needed by tests
|
# Export classes needed by tests
|
||||||
__all__ = [
|
__all__ = [
|
||||||
@@ -34,21 +25,9 @@ __all__ = [
|
|||||||
# Global manager instance
|
# Global manager instance
|
||||||
_manager = LLMSFullManager()
|
_manager = LLMSFullManager()
|
||||||
|
|
||||||
# Store root document first paragraph
|
|
||||||
_root_first_paragraph = ""
|
|
||||||
|
|
||||||
|
|
||||||
def doctree_resolved(app: Sphinx, doctree, docname: str):
|
def doctree_resolved(app: Sphinx, doctree, docname: str):
|
||||||
"""Called when a docname has been resolved to a document."""
|
"""Called when a docname has been resolved to a document."""
|
||||||
global _root_first_paragraph
|
|
||||||
|
|
||||||
# Check for llms-txt-ignore metadata at the page level
|
|
||||||
if hasattr(app.env, "metadata") and docname in app.env.metadata:
|
|
||||||
metadata = app.env.metadata[docname]
|
|
||||||
if metadata.get("llms-txt-ignore", "").lower() in ("true", "1", "yes"):
|
|
||||||
_manager.mark_page_ignored(docname)
|
|
||||||
return
|
|
||||||
|
|
||||||
# Extract title from the document
|
# Extract title from the document
|
||||||
title = None
|
title = None
|
||||||
# findall() returns a generator, convert to list to check if it has elements
|
# findall() returns a generator, convert to list to check if it has elements
|
||||||
@@ -59,14 +38,6 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
|
|||||||
if title:
|
if title:
|
||||||
_manager.update_page_title(docname, title)
|
_manager.update_page_title(docname, title)
|
||||||
|
|
||||||
# Extract first paragraph from root document
|
|
||||||
if docname == app.config.master_doc:
|
|
||||||
for node in doctree.traverse(nodes.paragraph):
|
|
||||||
first_para = node.astext()
|
|
||||||
if first_para:
|
|
||||||
_root_first_paragraph = first_para
|
|
||||||
break
|
|
||||||
|
|
||||||
|
|
||||||
def build_finished(app: Sphinx, exception):
|
def build_finished(app: Sphinx, exception):
|
||||||
"""Called when the build is finished."""
|
"""Called when the build is finished."""
|
||||||
@@ -76,25 +47,17 @@ def build_finished(app: Sphinx, exception):
|
|||||||
_manager.set_master_doc(app.config.master_doc)
|
_manager.set_master_doc(app.config.master_doc)
|
||||||
_manager.set_app(app)
|
_manager.set_app(app)
|
||||||
|
|
||||||
# Get the summary - use configured value or extracted first paragraph
|
|
||||||
summary = app.config.llms_txt_summary
|
|
||||||
if summary is None:
|
|
||||||
summary = _root_first_paragraph
|
|
||||||
|
|
||||||
# Set up configuration
|
# Set up configuration
|
||||||
config = {
|
config = {
|
||||||
"llms_txt_file": app.config.llms_txt_file,
|
"llms_txt_file": app.config.llms_txt_file,
|
||||||
"llms_txt_filename": app.config.llms_txt_filename,
|
"llms_txt_filename": app.config.llms_txt_filename,
|
||||||
"llms_txt_title": app.config.llms_txt_title,
|
"llms_txt_title": app.config.llms_txt_title,
|
||||||
"llms_txt_summary": summary,
|
"llms_txt_summary": app.config.llms_txt_summary,
|
||||||
"llms_txt_full_file": app.config.llms_txt_full_file,
|
"llms_txt_full_file": app.config.llms_txt_full_file,
|
||||||
"llms_txt_full_filename": app.config.llms_txt_full_filename,
|
"llms_txt_full_filename": app.config.llms_txt_full_filename,
|
||||||
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
||||||
"llms_txt_full_size_policy": app.config.llms_txt_full_size_policy,
|
|
||||||
"llms_txt_directives": app.config.llms_txt_directives,
|
"llms_txt_directives": app.config.llms_txt_directives,
|
||||||
"llms_txt_exclude": app.config.llms_txt_exclude,
|
"llms_txt_exclude": app.config.llms_txt_exclude,
|
||||||
"llms_txt_code_files": app.config.llms_txt_code_files,
|
|
||||||
"llms_txt_code_base_path": app.config.llms_txt_code_base_path,
|
|
||||||
"html_baseurl": getattr(app.config, "html_baseurl", ""),
|
"html_baseurl": getattr(app.config, "html_baseurl", ""),
|
||||||
}
|
}
|
||||||
_manager.set_config(config)
|
_manager.set_config(config)
|
||||||
@@ -113,33 +76,24 @@ def build_finished(app: Sphinx, exception):
|
|||||||
def setup(app: Sphinx) -> Dict[str, Any]:
|
def setup(app: Sphinx) -> Dict[str, Any]:
|
||||||
"""Set up the Sphinx extension."""
|
"""Set up the Sphinx extension."""
|
||||||
|
|
||||||
|
# Add configuration options
|
||||||
app.add_config_value("llms_txt_file", True, "env")
|
app.add_config_value("llms_txt_file", True, "env")
|
||||||
app.add_config_value("llms_txt_filename", "llms.txt", "env")
|
app.add_config_value("llms_txt_filename", "llms.txt", "env")
|
||||||
app.add_config_value("llms_txt_full_file", True, "env")
|
app.add_config_value("llms_txt_full_file", True, "env")
|
||||||
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
|
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
|
||||||
app.add_config_value("llms_txt_full_max_size", None, "env")
|
app.add_config_value("llms_txt_full_max_size", None, "env")
|
||||||
app.add_config_value("llms_txt_full_size_policy", "warn_skip", "env")
|
|
||||||
app.add_config_value("llms_txt_directives", [], "env")
|
app.add_config_value("llms_txt_directives", [], "env")
|
||||||
app.add_config_value("llms_txt_title", None, "env")
|
app.add_config_value("llms_txt_title", None, "env")
|
||||||
app.add_config_value("llms_txt_summary", None, "env")
|
app.add_config_value("llms_txt_summary", None, "env")
|
||||||
app.add_config_value("llms_txt_exclude", [], "env")
|
app.add_config_value("llms_txt_exclude", [], "env")
|
||||||
app.add_config_value("llms_txt_code_files", [], "env")
|
|
||||||
app.add_config_value("llms_txt_code_base_path", None, "env")
|
|
||||||
|
|
||||||
def builder_inited(app):
|
# Connect to Sphinx events
|
||||||
"""Used to limit what builders are allowed to run the extension."""
|
app.connect("doctree-resolved", doctree_resolved)
|
||||||
|
app.connect("build-finished", build_finished)
|
||||||
|
|
||||||
allowed_builders = ["html", "dirhtml"]
|
# Reset manager for each build
|
||||||
if hasattr(app, "builder") and app.builder.name in allowed_builders:
|
global _manager
|
||||||
# Reset manager and root paragraph for each build
|
_manager = LLMSFullManager()
|
||||||
global _manager, _root_first_paragraph
|
|
||||||
_manager = LLMSFullManager()
|
|
||||||
_root_first_paragraph = ""
|
|
||||||
|
|
||||||
app.connect("doctree-resolved", doctree_resolved)
|
|
||||||
app.connect("build-finished", build_finished)
|
|
||||||
|
|
||||||
app.connect("builder-inited", builder_inited)
|
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"version": __version__,
|
"version": __version__,
|
||||||
|
|||||||
+26
-125
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
import fnmatch
|
import fnmatch
|
||||||
from typing import Any, Dict, List, Tuple
|
from typing import Any, Dict, List
|
||||||
|
|
||||||
from sphinx.environment import BuildEnvironment
|
from sphinx.environment import BuildEnvironment
|
||||||
from sphinx.util import logging
|
from sphinx.util import logging
|
||||||
@@ -19,7 +19,6 @@ class DocumentCollector:
|
|||||||
self.master_doc: str = None
|
self.master_doc: str = None
|
||||||
self.env: BuildEnvironment = None
|
self.env: BuildEnvironment = None
|
||||||
self.config: Dict[str, Any] = {}
|
self.config: Dict[str, Any] = {}
|
||||||
self.app = None
|
|
||||||
|
|
||||||
def set_master_doc(self, master_doc: str):
|
def set_master_doc(self, master_doc: str):
|
||||||
"""Set the master document name."""
|
"""Set the master document name."""
|
||||||
@@ -38,79 +37,8 @@ class DocumentCollector:
|
|||||||
"""Set configuration options."""
|
"""Set configuration options."""
|
||||||
self.config = config
|
self.config = config
|
||||||
|
|
||||||
def set_app(self, app):
|
def get_page_order(self) -> List[str]:
|
||||||
"""Set the Sphinx application reference."""
|
"""Get the correct page order from the toctree structure."""
|
||||||
self.app = app
|
|
||||||
|
|
||||||
def _get_source_suffixes(self):
|
|
||||||
"""Get all valid source file suffixes from Sphinx configuration.
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
|
|
||||||
"""
|
|
||||||
if not self.app:
|
|
||||||
return [".rst"] # Default fallback
|
|
||||||
|
|
||||||
source_suffix = self.app.config.source_suffix
|
|
||||||
|
|
||||||
if isinstance(source_suffix, dict):
|
|
||||||
return list(source_suffix.keys())
|
|
||||||
elif isinstance(source_suffix, list):
|
|
||||||
return source_suffix
|
|
||||||
else:
|
|
||||||
return [source_suffix] # String format
|
|
||||||
|
|
||||||
def _get_docname_suffix(self, docname: str, sources_dir) -> str:
|
|
||||||
"""
|
|
||||||
Determine the source suffix for a given docname by checking which
|
|
||||||
file exists.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
docname: The document name to check
|
|
||||||
sources_dir: Path to the _sources directory
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
The source suffix if found, or None if no matching file exists
|
|
||||||
"""
|
|
||||||
if not sources_dir or not sources_dir.exists():
|
|
||||||
return None
|
|
||||||
|
|
||||||
# Get the source link suffix from Sphinx config
|
|
||||||
source_link_suffix = ""
|
|
||||||
if self.app and hasattr(self.app.config, "html_sourcelink_suffix"):
|
|
||||||
source_link_suffix = self.app.config.html_sourcelink_suffix
|
|
||||||
# Handle empty string case specially
|
|
||||||
if source_link_suffix == "":
|
|
||||||
source_link_suffix = "" # Keep it empty
|
|
||||||
elif not source_link_suffix.startswith("."):
|
|
||||||
source_link_suffix = "." + source_link_suffix
|
|
||||||
|
|
||||||
# Get the source file suffixes from Sphinx config
|
|
||||||
source_suffixes = self._get_source_suffixes()
|
|
||||||
|
|
||||||
# Try to find the source file with any of the valid source suffixes
|
|
||||||
for src_suffix in source_suffixes:
|
|
||||||
# Avoid duplicate extensions when source_suffix == source_link_suffix
|
|
||||||
if src_suffix == source_link_suffix:
|
|
||||||
candidate_file = sources_dir / f"{docname}{src_suffix}"
|
|
||||||
else:
|
|
||||||
candidate_file = (
|
|
||||||
sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
|
|
||||||
)
|
|
||||||
if candidate_file.exists():
|
|
||||||
return src_suffix
|
|
||||||
|
|
||||||
return None
|
|
||||||
|
|
||||||
def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
|
|
||||||
"""Get the correct page order from the toctree structure.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
sources_dir: Optional path to _sources directory for suffix detection
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
List of tuples (docname, source_suffix) in toctree order
|
|
||||||
"""
|
|
||||||
if not self.env or not self.master_doc:
|
if not self.env or not self.master_doc:
|
||||||
return []
|
return []
|
||||||
|
|
||||||
@@ -124,12 +52,9 @@ class DocumentCollector:
|
|||||||
|
|
||||||
visited.add(docname)
|
visited.add(docname)
|
||||||
|
|
||||||
# Add the current document with its suffix
|
# Add the current document
|
||||||
if docname not in [doc for doc, _ in page_order]:
|
if docname not in page_order:
|
||||||
suffix = None
|
page_order.append(docname)
|
||||||
if sources_dir:
|
|
||||||
suffix = self._get_docname_suffix(docname, sources_dir)
|
|
||||||
page_order.append((docname, suffix))
|
|
||||||
|
|
||||||
# Check for toctree entries in this document
|
# Check for toctree entries in this document
|
||||||
try:
|
try:
|
||||||
@@ -140,34 +65,21 @@ class DocumentCollector:
|
|||||||
):
|
):
|
||||||
for child_docname in self.env.toctree_includes[docname]:
|
for child_docname in self.env.toctree_includes[docname]:
|
||||||
collect_from_toctree(child_docname)
|
collect_from_toctree(child_docname)
|
||||||
# Try to use dependencies to find related documents
|
else:
|
||||||
elif (
|
# Fallback: try to resolve and parse the toctree
|
||||||
hasattr(self.env, "dependencies")
|
toctree = self.env.get_and_resolve_toctree(docname, None)
|
||||||
and docname in self.env.dependencies
|
if toctree:
|
||||||
):
|
from docutils import nodes
|
||||||
# Extract the dependent documents from the dependencies dict
|
|
||||||
for child_docname in self.env.dependencies[docname]:
|
|
||||||
# Only add documents actually in the document set
|
|
||||||
if (
|
|
||||||
hasattr(self.env, "all_docs")
|
|
||||||
and child_docname in self.env.all_docs
|
|
||||||
):
|
|
||||||
collect_from_toctree(child_docname)
|
|
||||||
# Fallback to titles or other available references
|
|
||||||
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
|
|
||||||
# Get all document names
|
|
||||||
all_docnames = list(self.env.all_docs.keys())
|
|
||||||
|
|
||||||
# Look for documents that might be related (have similar paths)
|
for node in list(toctree.findall(nodes.reference)):
|
||||||
current_prefix = "/".join(docname.split("/")[:-1])
|
if "refuri" in node.attributes:
|
||||||
if current_prefix:
|
refuri = node.attributes["refuri"]
|
||||||
for child_docname in all_docnames:
|
if refuri and refuri.endswith(".html"):
|
||||||
# Documents in the same directory might be related
|
child_docname = refuri[:-5] # Remove .html
|
||||||
if (
|
if (
|
||||||
child_docname.startswith(current_prefix)
|
child_docname != docname
|
||||||
and child_docname != docname
|
): # Avoid circular references
|
||||||
):
|
collect_from_toctree(child_docname)
|
||||||
collect_from_toctree(child_docname)
|
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.debug(f"Could not get toctree for {docname}: {e}")
|
logger.debug(f"Could not get toctree for {docname}: {e}")
|
||||||
|
|
||||||
@@ -176,33 +88,22 @@ class DocumentCollector:
|
|||||||
|
|
||||||
# Add any remaining documents not in the toctree (sorted)
|
# Add any remaining documents not in the toctree (sorted)
|
||||||
if hasattr(self.env, "all_docs"):
|
if hasattr(self.env, "all_docs"):
|
||||||
processed_docnames = {doc for doc, _ in page_order}
|
|
||||||
remaining = sorted(
|
remaining = sorted(
|
||||||
[
|
[doc for doc in self.env.all_docs.keys() if doc not in page_order]
|
||||||
doc
|
|
||||||
for doc in self.env.all_docs.keys()
|
|
||||||
if doc not in processed_docnames
|
|
||||||
]
|
|
||||||
)
|
)
|
||||||
for docname in remaining:
|
page_order.extend(remaining)
|
||||||
suffix = None
|
|
||||||
if sources_dir:
|
|
||||||
suffix = self._get_docname_suffix(docname, sources_dir)
|
|
||||||
page_order.append((docname, suffix))
|
|
||||||
|
|
||||||
return page_order
|
return page_order
|
||||||
|
|
||||||
def filter_excluded_pages(
|
def filter_excluded_pages(self, page_order: List[str]) -> List[str]:
|
||||||
self, page_order: List[Tuple[str, str]]
|
|
||||||
) -> List[Tuple[str, str]]:
|
|
||||||
"""Filter out excluded pages from the page order."""
|
"""Filter out excluded pages from the page order."""
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
exclude_patterns = self.config.get("llms_txt_exclude")
|
||||||
if exclude_patterns:
|
if exclude_patterns:
|
||||||
return [
|
return [
|
||||||
(docname, suffix)
|
page
|
||||||
for docname, suffix in page_order
|
for page in page_order
|
||||||
if not any(
|
if not any(
|
||||||
self._match_exclude_pattern(docname, pattern)
|
self._match_exclude_pattern(page, pattern)
|
||||||
for pattern in exclude_patterns
|
for pattern in exclude_patterns
|
||||||
)
|
)
|
||||||
]
|
]
|
||||||
|
|||||||
+128
-779
File diff suppressed because it is too large
Load Diff
+41
-143
@@ -44,10 +44,7 @@ class DocumentProcessor:
|
|||||||
Returns:
|
Returns:
|
||||||
Processed content with directives properly resolved
|
Processed content with directives properly resolved
|
||||||
"""
|
"""
|
||||||
# First process llms-txt-ignore blocks
|
# First process include directives
|
||||||
content = self._process_ignore_blocks(content)
|
|
||||||
|
|
||||||
# Then process include directives
|
|
||||||
content = self._process_includes(content, source_path)
|
content = self._process_includes(content, source_path)
|
||||||
|
|
||||||
# Then process path directives (image, figure, etc.)
|
# Then process path directives (image, figure, etc.)
|
||||||
@@ -97,14 +94,8 @@ class DocumentProcessor:
|
|||||||
if not base_url:
|
if not base_url:
|
||||||
return path
|
return path
|
||||||
|
|
||||||
# Ensure base URL ends with slash
|
|
||||||
if not base_url.endswith("/"):
|
if not base_url.endswith("/"):
|
||||||
base_url += "/"
|
base_url += "/"
|
||||||
|
|
||||||
# Remove leading slash from path to avoid double slashes
|
|
||||||
if path.startswith("/"):
|
|
||||||
path = path[1:]
|
|
||||||
|
|
||||||
return f"{base_url}{path}"
|
return f"{base_url}{path}"
|
||||||
|
|
||||||
def _is_absolute_or_url(self, path: str) -> bool:
|
def _is_absolute_or_url(self, path: str) -> bool:
|
||||||
@@ -146,73 +137,12 @@ class DocumentProcessor:
|
|||||||
prefix = match.group(1) # The entire directive prefix including whitespace
|
prefix = match.group(1) # The entire directive prefix including whitespace
|
||||||
path = match.group(3).strip() # The path argument
|
path = match.group(3).strip() # The path argument
|
||||||
|
|
||||||
# Handle URLs and data URIs - leave unchanged
|
# Only process relative paths, not absolute paths or URLs
|
||||||
if path.startswith(("http://", "https://", "data:")):
|
if not self._is_absolute_or_url(path):
|
||||||
return match.group(0)
|
# Special case for test files
|
||||||
|
if is_test:
|
||||||
# For ALL paths, check if image exists in _images first
|
# Add subdir/ prefix to match test expectations
|
||||||
# Extract filename from the path
|
full_path = "subdir/" + path
|
||||||
filename = os.path.basename(path)
|
|
||||||
|
|
||||||
# Check if image exists in _images directory
|
|
||||||
# First determine the build directory from source_path
|
|
||||||
build_dir = None
|
|
||||||
if "_sources" in str(source_path):
|
|
||||||
# Extract build directory (parent of _sources)
|
|
||||||
path_parts = str(source_path).split("_sources/")
|
|
||||||
if len(path_parts) > 1:
|
|
||||||
build_dir = path_parts[0].rstrip("/")
|
|
||||||
|
|
||||||
# If we can determine the build directory, check if image exists in _images
|
|
||||||
if build_dir:
|
|
||||||
images_path = os.path.join(build_dir, "_images", filename)
|
|
||||||
if os.path.exists(images_path):
|
|
||||||
# Image exists in _images, use _images path
|
|
||||||
full_path = f"/_images/{filename}"
|
|
||||||
# Add base URL if configured
|
|
||||||
full_path = self._add_base_url(full_path, base_url)
|
|
||||||
return f"{prefix}{full_path}"
|
|
||||||
|
|
||||||
# Image doesn't exist in _images, handle based on path type
|
|
||||||
# Handle absolute paths (starting with /) - add base URL if configured
|
|
||||||
if path.startswith("/"):
|
|
||||||
# Add base URL to absolute paths if configured
|
|
||||||
full_path = self._add_base_url(path, base_url)
|
|
||||||
return f"{prefix}{full_path}"
|
|
||||||
|
|
||||||
# Handle relative paths with original logic for backward compatibility
|
|
||||||
# Special case for test files
|
|
||||||
if is_test:
|
|
||||||
# Add subdir/ prefix to match test expectations
|
|
||||||
full_path = "subdir/" + path
|
|
||||||
|
|
||||||
# If base_url is set, prepend it to the path
|
|
||||||
full_path = self._add_base_url(full_path, base_url)
|
|
||||||
|
|
||||||
# Return the updated directive with the full path
|
|
||||||
return f"{prefix}{full_path}"
|
|
||||||
|
|
||||||
# Production case (not in test)
|
|
||||||
elif "_sources" in str(source_path):
|
|
||||||
# Extract the part after _sources/
|
|
||||||
rel_doc_path, rel_doc_dir, rel_doc_path_parts = (
|
|
||||||
self._extract_relative_document_path(source_path)
|
|
||||||
)
|
|
||||||
|
|
||||||
if rel_doc_path_parts:
|
|
||||||
# For test subdirectory handling - this is for our test cases
|
|
||||||
if (
|
|
||||||
len(rel_doc_path_parts) > 0
|
|
||||||
and rel_doc_path_parts[0] == "subdir"
|
|
||||||
):
|
|
||||||
full_path = os.path.normpath(os.path.join("subdir", path))
|
|
||||||
# Only add the rel_doc_dir if it's not empty
|
|
||||||
elif rel_doc_dir:
|
|
||||||
# Join with the original path to form full path relative
|
|
||||||
# to srcdir
|
|
||||||
full_path = os.path.normpath(os.path.join(rel_doc_dir, path))
|
|
||||||
else:
|
|
||||||
full_path = path
|
|
||||||
|
|
||||||
# If base_url is set, prepend it to the path
|
# If base_url is set, prepend it to the path
|
||||||
full_path = self._add_base_url(full_path, base_url)
|
full_path = self._add_base_url(full_path, base_url)
|
||||||
@@ -220,12 +150,37 @@ class DocumentProcessor:
|
|||||||
# Return the updated directive with the full path
|
# Return the updated directive with the full path
|
||||||
return f"{prefix}{full_path}"
|
return f"{prefix}{full_path}"
|
||||||
|
|
||||||
# Fallback for relative paths - add base URL if configured
|
# Production case (not in test)
|
||||||
else:
|
elif "_sources" in str(source_path):
|
||||||
full_path = self._add_base_url(path, base_url)
|
# Extract the part after _sources/
|
||||||
return f"{prefix}{full_path}"
|
rel_doc_path, rel_doc_dir, rel_doc_path_parts = (
|
||||||
|
self._extract_relative_document_path(source_path)
|
||||||
|
)
|
||||||
|
|
||||||
# If we couldn't resolve the path, return unchanged
|
if rel_doc_path_parts:
|
||||||
|
# For test subdirectory handling - this is for our test cases
|
||||||
|
if (
|
||||||
|
len(rel_doc_path_parts) > 0
|
||||||
|
and rel_doc_path_parts[0] == "subdir"
|
||||||
|
):
|
||||||
|
full_path = os.path.normpath(os.path.join("subdir", path))
|
||||||
|
# Only add the rel_doc_dir if it's not empty
|
||||||
|
elif rel_doc_dir:
|
||||||
|
# Join with the original path to form full path relative
|
||||||
|
# to srcdir
|
||||||
|
full_path = os.path.normpath(
|
||||||
|
os.path.join(rel_doc_dir, path)
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
full_path = path
|
||||||
|
|
||||||
|
# If base_url is set, prepend it to the path
|
||||||
|
full_path = self._add_base_url(full_path, base_url)
|
||||||
|
|
||||||
|
# Return the updated directive with the full path
|
||||||
|
return f"{prefix}{full_path}"
|
||||||
|
|
||||||
|
# If we couldn't resolve the path or it's already absolute, return unchanged
|
||||||
return match.group(0)
|
return match.group(0)
|
||||||
|
|
||||||
# Replace directive paths in the content
|
# Replace directive paths in the content
|
||||||
@@ -246,12 +201,9 @@ class DocumentProcessor:
|
|||||||
"""
|
"""
|
||||||
possible_paths = []
|
possible_paths = []
|
||||||
|
|
||||||
# If it's an absolute path, treat it as relative to srcdir
|
# If it's an absolute path, use it directly
|
||||||
if os.path.isabs(include_path):
|
if os.path.isabs(include_path):
|
||||||
# Remove the leading slash and treat as relative to srcdir
|
possible_paths.append(Path(include_path))
|
||||||
relative_path = include_path.lstrip("/")
|
|
||||||
if self.srcdir:
|
|
||||||
possible_paths.append((Path(self.srcdir) / relative_path).resolve())
|
|
||||||
else:
|
else:
|
||||||
# Relative to the source file (in _sources directory)
|
# Relative to the source file (in _sources directory)
|
||||||
possible_paths.append((source_path.parent / include_path).resolve())
|
possible_paths.append((source_path.parent / include_path).resolve())
|
||||||
@@ -292,9 +244,6 @@ class DocumentProcessor:
|
|||||||
# Function to replace each include with content
|
# Function to replace each include with content
|
||||||
def replace_include(match):
|
def replace_include(match):
|
||||||
include_path = match.group(3)
|
include_path = match.group(3)
|
||||||
directive_part = match.group(
|
|
||||||
1
|
|
||||||
) # The ".. include:: " part with leading whitespace
|
|
||||||
|
|
||||||
# Get all possible paths to try
|
# Get all possible paths to try
|
||||||
possible_paths = self._resolve_include_paths(include_path, source_path)
|
possible_paths = self._resolve_include_paths(include_path, source_path)
|
||||||
@@ -305,18 +254,7 @@ class DocumentProcessor:
|
|||||||
if path_to_try.exists():
|
if path_to_try.exists():
|
||||||
with open(path_to_try, "r", encoding="utf-8") as f:
|
with open(path_to_try, "r", encoding="utf-8") as f:
|
||||||
included_content = f.read()
|
included_content = f.read()
|
||||||
|
return included_content
|
||||||
# Find where the actual directive starts, after any whitespace
|
|
||||||
directive_start = directive_part.find("..")
|
|
||||||
if directive_start > 0:
|
|
||||||
# There's leading whitespace/newlines before the directive
|
|
||||||
leading_part = directive_part[:directive_start]
|
|
||||||
# Replace directive with content, preserving the structure
|
|
||||||
return leading_part + included_content
|
|
||||||
else:
|
|
||||||
# No leading whitespace, just return the content
|
|
||||||
return included_content
|
|
||||||
|
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.error(
|
logger.error(
|
||||||
f"sphinx-llms-txt: Error reading include file {path_to_try}:"
|
f"sphinx-llms-txt: Error reading include file {path_to_try}:"
|
||||||
@@ -328,48 +266,8 @@ class DocumentProcessor:
|
|||||||
paths_tried = ", ".join(str(p) for p in possible_paths)
|
paths_tried = ", ".join(str(p) for p in possible_paths)
|
||||||
logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}")
|
logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}")
|
||||||
logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}")
|
logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}")
|
||||||
|
return f"[Include file not found: {include_path}]"
|
||||||
# Preserve spacing structure for error message too
|
|
||||||
directive_start = match.group(1).find("..")
|
|
||||||
if directive_start > 0:
|
|
||||||
leading_part = match.group(1)[:directive_start]
|
|
||||||
return leading_part + f"[Include file not found: {include_path}]"
|
|
||||||
else:
|
|
||||||
return f"[Include file not found: {include_path}]"
|
|
||||||
|
|
||||||
# Replace all includes with their content
|
# Replace all includes with their content
|
||||||
processed_content = include_pattern.sub(replace_include, content)
|
processed_content = include_pattern.sub(replace_include, content)
|
||||||
return processed_content
|
return processed_content
|
||||||
|
|
||||||
def _process_ignore_blocks(self, content: str) -> str:
|
|
||||||
"""Process llms-txt-ignore-start/end blocks by removing their content.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
content: The source content to process
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Processed content with ignore blocks removed
|
|
||||||
"""
|
|
||||||
# Process ignore blocks iteratively to handle nested cases correctly
|
|
||||||
while True:
|
|
||||||
# Pattern to match ignore blocks - handles whitespace and indentation
|
|
||||||
ignore_pattern = re.compile(
|
|
||||||
r"^\s*\.\.\s+llms-txt-ignore-start\s*\n" # Start directive line
|
|
||||||
r"(.*?)" # Content to ignore (non-greedy)
|
|
||||||
r"^\s*\.\.\s+llms-txt-ignore-end\s*$", # End directive line
|
|
||||||
re.MULTILINE | re.DOTALL,
|
|
||||||
)
|
|
||||||
|
|
||||||
# Find and remove one ignore block at a time
|
|
||||||
match = ignore_pattern.search(content)
|
|
||||||
if not match:
|
|
||||||
break
|
|
||||||
|
|
||||||
# Remove the matched block
|
|
||||||
content = content[: match.start()] + content[match.end() :]
|
|
||||||
|
|
||||||
# Clean up any extra blank lines that might be left
|
|
||||||
# Replace multiple consecutive newlines with at most 2 newlines
|
|
||||||
processed_content = re.sub(r"\n\n\n+", "\n\n", content)
|
|
||||||
|
|
||||||
return processed_content
|
|
||||||
|
|||||||
+10
-17
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Dict, List, Tuple, Union
|
from typing import Any, Dict, List
|
||||||
|
|
||||||
from sphinx.application import Sphinx
|
from sphinx.application import Sphinx
|
||||||
from sphinx.util import logging
|
from sphinx.util import logging
|
||||||
@@ -37,24 +37,24 @@ class FileWriter:
|
|||||||
f.write("\n".join(content_parts))
|
f.write("\n".join(content_parts))
|
||||||
|
|
||||||
logger.info(
|
logger.info(
|
||||||
f"sphinx-llms-txt: Created {output_path} with {len(content_parts)}"
|
f"sphinx-llms-txt: created {output_path} with {len(content_parts)}"
|
||||||
f" sources and {total_line_count} lines"
|
f" sources and {total_line_count} lines"
|
||||||
)
|
)
|
||||||
return True
|
return True
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.error(f"sphinx-llms-txt: Error writing combined sources file: {e}")
|
logger.error(f"sphinx-llm-txt: Error writing combined sources file: {e}")
|
||||||
return False
|
return False
|
||||||
|
|
||||||
def write_verbose_info_to_file(
|
def write_verbose_info_to_file(
|
||||||
self,
|
self,
|
||||||
page_order: Union[List[str], List[Tuple[str, str]]],
|
page_order: List[str],
|
||||||
page_titles: Dict[str, str],
|
page_titles: Dict[str, str],
|
||||||
total_line_count: int = 0,
|
total_line_count: int = 0,
|
||||||
) -> bool:
|
) -> bool:
|
||||||
"""Write summary information to the llms.txt file.
|
"""Write summary information to the llms.txt file.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
page_order: Ordered list of document names or (docname, suffix) tuples
|
page_order: Ordered list of document names
|
||||||
page_titles: Dictionary mapping docnames to titles
|
page_titles: Dictionary mapping docnames to titles
|
||||||
total_line_count: Total number of lines in the combined content
|
total_line_count: Total number of lines in the combined content
|
||||||
|
|
||||||
@@ -88,12 +88,10 @@ class FileWriter:
|
|||||||
if description:
|
if description:
|
||||||
# Trim leading and trailing whitespace
|
# Trim leading and trailing whitespace
|
||||||
description = description.strip()
|
description = description.strip()
|
||||||
if description:
|
# Replace newlines with newline + blockquote marker to maintain
|
||||||
# Only add blockquote if description is not empty
|
# blockquote formatting
|
||||||
# Replace newlines with newline + blockquote marker to maintain
|
description = description.replace("\n", "\n> ")
|
||||||
# blockquote formatting
|
f.write(f"> {description}\n\n")
|
||||||
description = description.replace("\n", "\n> ")
|
|
||||||
f.write(f"> {description}\n\n")
|
|
||||||
|
|
||||||
f.write("## Docs\n\n")
|
f.write("## Docs\n\n")
|
||||||
# Get base URL from config
|
# Get base URL from config
|
||||||
@@ -102,12 +100,7 @@ class FileWriter:
|
|||||||
if not base_url.endswith("/"):
|
if not base_url.endswith("/"):
|
||||||
base_url += "/"
|
base_url += "/"
|
||||||
|
|
||||||
for item in page_order:
|
for docname in page_order:
|
||||||
# Handle both old format (str) and new format (tuple)
|
|
||||||
if isinstance(item, tuple):
|
|
||||||
docname, _ = item
|
|
||||||
else:
|
|
||||||
docname = item
|
|
||||||
title = page_titles.get(docname, docname)
|
title = page_titles.get(docname, docname)
|
||||||
f.write(f"- [{title}]({base_url}{docname}.html)\n")
|
f.write(f"- [{title}]({base_url}{docname}.html)\n")
|
||||||
|
|
||||||
|
|||||||
@@ -8,8 +8,6 @@ Welcome to Test Project's documentation!
|
|||||||
page1
|
page1
|
||||||
page2
|
page2
|
||||||
page_with_include
|
page_with_include
|
||||||
page_ignored_metadata
|
|
||||||
page_with_ignore_blocks
|
|
||||||
|
|
||||||
Indices and tables
|
Indices and tables
|
||||||
==================
|
==================
|
||||||
|
|||||||
@@ -1,16 +0,0 @@
|
|||||||
:llms-txt-ignore: true
|
|
||||||
|
|
||||||
Page Ignored by Metadata
|
|
||||||
========================
|
|
||||||
|
|
||||||
This page should not appear in llms-full.txt because of the metadata directive.
|
|
||||||
|
|
||||||
Section 1
|
|
||||||
---------
|
|
||||||
|
|
||||||
This content should be completely ignored.
|
|
||||||
|
|
||||||
Section 2
|
|
||||||
---------
|
|
||||||
|
|
||||||
This content should also be ignored.
|
|
||||||
@@ -1,39 +0,0 @@
|
|||||||
Page With Ignore Blocks
|
|
||||||
=======================
|
|
||||||
|
|
||||||
This content should appear in llms-full.txt.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
This content should be ignored and not appear in llms-full.txt.
|
|
||||||
|
|
||||||
Section Ignored
|
|
||||||
---------------
|
|
||||||
|
|
||||||
This section should also be ignored.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
This content after the ignore block should appear in llms-full.txt.
|
|
||||||
|
|
||||||
Another Section
|
|
||||||
---------------
|
|
||||||
|
|
||||||
This content should definitely appear.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
Another ignored block with multiple lines.
|
|
||||||
|
|
||||||
- Item 1 (ignored)
|
|
||||||
- Item 2 (ignored)
|
|
||||||
|
|
||||||
.. code-block:: python
|
|
||||||
|
|
||||||
# This code should be ignored
|
|
||||||
def ignored_function():
|
|
||||||
pass
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
Final content that should appear.
|
|
||||||
@@ -1,220 +0,0 @@
|
|||||||
"""Tests for llms-txt ignore features."""
|
|
||||||
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
from sphinx_llms_txt import DocumentProcessor
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_ignore_blocks():
|
|
||||||
"""Test that ignore blocks are properly removed from content."""
|
|
||||||
processor = DocumentProcessor({}, None)
|
|
||||||
|
|
||||||
content = """This content should remain.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
This content should be removed.
|
|
||||||
|
|
||||||
Section Ignored
|
|
||||||
---------------
|
|
||||||
|
|
||||||
This section should also be removed.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
This content should remain after the ignore block.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
Another ignored block.
|
|
||||||
Multiple lines here.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
Final content that should remain."""
|
|
||||||
|
|
||||||
processed = processor._process_ignore_blocks(content)
|
|
||||||
|
|
||||||
# Check that ignored content is removed
|
|
||||||
assert "This content should be removed." not in processed
|
|
||||||
assert "Section Ignored" not in processed
|
|
||||||
assert "Another ignored block." not in processed
|
|
||||||
assert "Multiple lines here." not in processed
|
|
||||||
|
|
||||||
# Check that non-ignored content remains
|
|
||||||
assert "This content should remain." in processed
|
|
||||||
assert "This content should remain after the ignore block." in processed
|
|
||||||
assert "Final content that should remain." in processed
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_ignore_blocks_with_indentation():
|
|
||||||
"""Test that ignore blocks work with different indentation levels."""
|
|
||||||
processor = DocumentProcessor({}, None)
|
|
||||||
|
|
||||||
content = """Section Title
|
|
||||||
=============
|
|
||||||
|
|
||||||
Normal content.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
Indented ignored content.
|
|
||||||
More indented content.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
Back to normal content."""
|
|
||||||
|
|
||||||
processed = processor._process_ignore_blocks(content)
|
|
||||||
|
|
||||||
# Check that ignored content is removed
|
|
||||||
assert "Indented ignored content." not in processed
|
|
||||||
assert "More indented content." not in processed
|
|
||||||
|
|
||||||
# Check that non-ignored content remains
|
|
||||||
assert "Section Title" in processed
|
|
||||||
assert "Normal content." in processed
|
|
||||||
assert "Back to normal content." in processed
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_ignore_blocks_multiple():
|
|
||||||
"""Test that multiple ignore blocks are handled correctly."""
|
|
||||||
processor = DocumentProcessor({}, None)
|
|
||||||
|
|
||||||
content = """Start content.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
First ignore block.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
Middle content that should remain.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
Second ignore block.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
End content."""
|
|
||||||
|
|
||||||
processed = processor._process_ignore_blocks(content)
|
|
||||||
|
|
||||||
# Check that ignored content is removed
|
|
||||||
assert "First ignore block." not in processed
|
|
||||||
assert "Second ignore block." not in processed
|
|
||||||
|
|
||||||
# Check that non-ignored content remains
|
|
||||||
assert "Start content." in processed
|
|
||||||
assert "Middle content that should remain." in processed
|
|
||||||
assert "End content." in processed
|
|
||||||
|
|
||||||
|
|
||||||
def test_build_with_ignore_features(basic_sphinx_app):
|
|
||||||
"""Test building HTML documentation with ignore features."""
|
|
||||||
app = basic_sphinx_app
|
|
||||||
app.build()
|
|
||||||
|
|
||||||
# Check if the output file was created
|
|
||||||
output_file = Path(app.outdir) / "test-llms-full.txt"
|
|
||||||
assert output_file.exists(), f"Output file {output_file} does not exist"
|
|
||||||
|
|
||||||
# Read the content of the output file
|
|
||||||
content = output_file.read_text()
|
|
||||||
|
|
||||||
# Check that page with metadata ignore is completely excluded
|
|
||||||
assert "Page Ignored by Metadata" not in content
|
|
||||||
assert "This page should not appear in llms-full.txt" not in content
|
|
||||||
|
|
||||||
# Check that page with ignore blocks has the right content
|
|
||||||
assert "Page With Ignore Blocks" in content
|
|
||||||
assert "This content should appear in llms-full.txt." in content
|
|
||||||
assert "This content after the ignore block should appear" in content
|
|
||||||
assert "Another Section" in content
|
|
||||||
assert "Final content that should appear." in content
|
|
||||||
|
|
||||||
# Check that ignored block content is not present
|
|
||||||
assert "This content should be ignored and not appear" not in content
|
|
||||||
assert "Section Ignored" not in content
|
|
||||||
assert "Another ignored block with multiple lines." not in content
|
|
||||||
assert "Item 1 (ignored)" not in content
|
|
||||||
assert "def ignored_function():" not in content
|
|
||||||
|
|
||||||
|
|
||||||
def test_manager_mark_page_ignored():
|
|
||||||
"""Test that manager can mark pages as ignored."""
|
|
||||||
from sphinx_llms_txt import LLMSFullManager
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
|
|
||||||
# Initially no pages are ignored
|
|
||||||
assert len(manager.ignored_pages) == 0
|
|
||||||
|
|
||||||
# Mark a page as ignored
|
|
||||||
manager.mark_page_ignored("test_page")
|
|
||||||
|
|
||||||
# Check that page is in ignored set
|
|
||||||
assert "test_page" in manager.ignored_pages
|
|
||||||
assert len(manager.ignored_pages) == 1
|
|
||||||
|
|
||||||
# Mark another page as ignored
|
|
||||||
manager.mark_page_ignored("another_page")
|
|
||||||
|
|
||||||
# Check both pages are ignored
|
|
||||||
assert "test_page" in manager.ignored_pages
|
|
||||||
assert "another_page" in manager.ignored_pages
|
|
||||||
assert len(manager.ignored_pages) == 2
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_ignore_blocks_empty_blocks():
|
|
||||||
"""Test that empty ignore blocks are handled correctly."""
|
|
||||||
processor = DocumentProcessor({}, None)
|
|
||||||
|
|
||||||
content = """Content before.
|
|
||||||
|
|
||||||
.. llms-txt-ignore-start
|
|
||||||
|
|
||||||
.. llms-txt-ignore-end
|
|
||||||
|
|
||||||
Content after."""
|
|
||||||
|
|
||||||
processed = processor._process_ignore_blocks(content)
|
|
||||||
|
|
||||||
# Check that content remains
|
|
||||||
assert "Content before." in processed
|
|
||||||
assert "Content after." in processed
|
|
||||||
|
|
||||||
# Check that we don't have excessive newlines
|
|
||||||
lines = processed.strip().split("\n")
|
|
||||||
non_empty_lines = [line for line in lines if line.strip()]
|
|
||||||
assert len(non_empty_lines) == 2
|
|
||||||
|
|
||||||
|
|
||||||
def test_ignore_metadata_affects_both_files(basic_sphinx_app):
|
|
||||||
"""Test that :llms-txt-ignore: true affects both files."""
|
|
||||||
app = basic_sphinx_app
|
|
||||||
# Enable both llms.txt and llms-full.txt file generation
|
|
||||||
app.config.llms_txt_file = True
|
|
||||||
app.config.llms_txt_filename = "test-llms.txt"
|
|
||||||
app.build()
|
|
||||||
|
|
||||||
# Check if both output files were created
|
|
||||||
llms_full_file = Path(app.outdir) / "test-llms-full.txt"
|
|
||||||
llms_summary_file = Path(app.outdir) / "test-llms.txt"
|
|
||||||
|
|
||||||
assert llms_full_file.exists(), f"Output file {llms_full_file} does not exist"
|
|
||||||
assert llms_summary_file.exists(), f"Output file {llms_summary_file} does not exist"
|
|
||||||
|
|
||||||
# Read the content of both files
|
|
||||||
llms_full_content = llms_full_file.read_text()
|
|
||||||
llms_summary_content = llms_summary_file.read_text()
|
|
||||||
|
|
||||||
# Check that page with metadata ignore is excluded from llms-full.txt
|
|
||||||
assert "Page Ignored by Metadata" not in llms_full_content
|
|
||||||
assert "This page should not appear in llms-full.txt" not in llms_full_content
|
|
||||||
|
|
||||||
# Check that page with metadata ignore is also excluded from llms.txt
|
|
||||||
# This should NOT contain a link to the ignored page
|
|
||||||
assert "Page Ignored by Metadata" not in llms_summary_content
|
|
||||||
assert "page_ignored_metadata.html" not in llms_summary_content
|
|
||||||
@@ -117,152 +117,6 @@ def test_max_lines_limit(temp_dir, rootdir):
|
|||||||
app.docutils_conf_path.unlink()
|
app.docutils_conf_path.unlink()
|
||||||
|
|
||||||
|
|
||||||
def test_on_exceed_skip(temp_dir, rootdir):
|
|
||||||
"""Test that skip action works when size limit is exceeded."""
|
|
||||||
from sphinx.testing.util import SphinxTestApp
|
|
||||||
|
|
||||||
src_dir = rootdir / "basic"
|
|
||||||
|
|
||||||
app = SphinxTestApp(
|
|
||||||
srcdir=src_dir,
|
|
||||||
builddir=temp_dir,
|
|
||||||
buildername="html",
|
|
||||||
freshenv=True,
|
|
||||||
confoverrides={
|
|
||||||
"llms_txt_full_filename": "skip-test.txt",
|
|
||||||
"llms_txt_full_max_size": 20,
|
|
||||||
"llms_txt_full_size_policy": "warn_skip",
|
|
||||||
},
|
|
||||||
)
|
|
||||||
|
|
||||||
app.build()
|
|
||||||
|
|
||||||
# Check that the output file was NOT created
|
|
||||||
output_file = Path(app.outdir) / "skip-test.txt"
|
|
||||||
assert (
|
|
||||||
not output_file.exists()
|
|
||||||
), f"Output file {output_file} should not exist with skip action"
|
|
||||||
|
|
||||||
# Cleanup
|
|
||||||
sys.path[:] = app._saved_path
|
|
||||||
_clean_up_global_state()
|
|
||||||
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
|
||||||
app.docutils_conf_path.unlink()
|
|
||||||
|
|
||||||
|
|
||||||
def test_on_exceed_keep(temp_dir, rootdir):
|
|
||||||
"""Test that keep action works when size limit is exceeded."""
|
|
||||||
from sphinx.testing.util import SphinxTestApp
|
|
||||||
|
|
||||||
src_dir = rootdir / "basic"
|
|
||||||
|
|
||||||
app = SphinxTestApp(
|
|
||||||
srcdir=src_dir,
|
|
||||||
builddir=temp_dir,
|
|
||||||
buildername="html",
|
|
||||||
freshenv=True,
|
|
||||||
confoverrides={
|
|
||||||
"llms_txt_full_filename": "keep-test.txt",
|
|
||||||
"llms_txt_full_max_size": 20,
|
|
||||||
"llms_txt_full_size_policy": "info_keep",
|
|
||||||
},
|
|
||||||
)
|
|
||||||
|
|
||||||
app.build()
|
|
||||||
|
|
||||||
# Check that the output file WAS created despite exceeding limit
|
|
||||||
output_file = Path(app.outdir) / "keep-test.txt"
|
|
||||||
assert (
|
|
||||||
output_file.exists()
|
|
||||||
), f"Output file {output_file} should exist with keep action"
|
|
||||||
|
|
||||||
# Verify it has content
|
|
||||||
content = output_file.read_text()
|
|
||||||
assert len(content) > 0, "Output file should have content with keep action"
|
|
||||||
|
|
||||||
# Cleanup
|
|
||||||
sys.path[:] = app._saved_path
|
|
||||||
_clean_up_global_state()
|
|
||||||
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
|
||||||
app.docutils_conf_path.unlink()
|
|
||||||
|
|
||||||
|
|
||||||
def test_on_exceed_note(temp_dir, rootdir):
|
|
||||||
"""Test that note action works when size limit is exceeded."""
|
|
||||||
from sphinx.testing.util import SphinxTestApp
|
|
||||||
|
|
||||||
src_dir = rootdir / "basic"
|
|
||||||
|
|
||||||
app = SphinxTestApp(
|
|
||||||
srcdir=src_dir,
|
|
||||||
builddir=temp_dir,
|
|
||||||
buildername="html",
|
|
||||||
freshenv=True,
|
|
||||||
confoverrides={
|
|
||||||
"llms_txt_full_filename": "note-test.txt",
|
|
||||||
"llms_txt_full_max_size": 20,
|
|
||||||
"llms_txt_full_size_policy": "warn_note",
|
|
||||||
},
|
|
||||||
)
|
|
||||||
|
|
||||||
app.build()
|
|
||||||
|
|
||||||
# Check that the output file WAS created with placeholder content
|
|
||||||
output_file = Path(app.outdir) / "note-test.txt"
|
|
||||||
assert (
|
|
||||||
output_file.exists()
|
|
||||||
), f"Output file {output_file} should exist with note action"
|
|
||||||
|
|
||||||
# Verify it has the placeholder content
|
|
||||||
content = output_file.read_text()
|
|
||||||
assert (
|
|
||||||
"This file was not generated because it exceeded the configured size limit."
|
|
||||||
in content
|
|
||||||
)
|
|
||||||
assert "llms_txt_full_max_size" in content
|
|
||||||
assert "llms_txt_full_size_policy" in content
|
|
||||||
assert "Configured max size: 20 lines" in content
|
|
||||||
|
|
||||||
# Cleanup
|
|
||||||
sys.path[:] = app._saved_path
|
|
||||||
_clean_up_global_state()
|
|
||||||
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
|
||||||
app.docutils_conf_path.unlink()
|
|
||||||
|
|
||||||
|
|
||||||
def test_on_exceed_invalid_config(temp_dir, rootdir):
|
|
||||||
"""Test behavior with invalid configuration values."""
|
|
||||||
from sphinx.testing.util import SphinxTestApp
|
|
||||||
|
|
||||||
src_dir = rootdir / "basic"
|
|
||||||
|
|
||||||
app = SphinxTestApp(
|
|
||||||
srcdir=src_dir,
|
|
||||||
builddir=temp_dir,
|
|
||||||
buildername="html",
|
|
||||||
freshenv=True,
|
|
||||||
confoverrides={
|
|
||||||
"llms_txt_full_filename": "invalid-test.txt",
|
|
||||||
"llms_txt_full_max_size": 20,
|
|
||||||
"llms_txt_full_size_policy": "invalid_format", # Invalid config
|
|
||||||
},
|
|
||||||
)
|
|
||||||
|
|
||||||
app.build()
|
|
||||||
|
|
||||||
# Should fall back to default behavior (warn_skip)
|
|
||||||
output_file = Path(app.outdir) / "invalid-test.txt"
|
|
||||||
assert (
|
|
||||||
not output_file.exists()
|
|
||||||
), f"Output file {output_file} should not exist with invalid config fallback"
|
|
||||||
|
|
||||||
# Cleanup
|
|
||||||
sys.path[:] = app._saved_path
|
|
||||||
_clean_up_global_state()
|
|
||||||
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
|
||||||
app.docutils_conf_path.unlink()
|
|
||||||
|
|
||||||
|
|
||||||
def test_title_override(temp_dir, rootdir):
|
def test_title_override(temp_dir, rootdir):
|
||||||
"""Test that the title override works correctly."""
|
"""Test that the title override works correctly."""
|
||||||
from sphinx.testing.util import SphinxTestApp
|
from sphinx.testing.util import SphinxTestApp
|
||||||
|
|||||||
@@ -41,42 +41,6 @@ def test_setup_returns_valid_dict():
|
|||||||
assert "parallel_write_safe" in result
|
assert "parallel_write_safe" in result
|
||||||
|
|
||||||
|
|
||||||
def test_builder_inited_with_disallowed_builder():
|
|
||||||
"""Test that disallowed builders do not trigger extension setup."""
|
|
||||||
import sphinx_llms_txt
|
|
||||||
|
|
||||||
# Reset global state
|
|
||||||
sphinx_llms_txt._manager = sphinx_llms_txt.LLMSFullManager()
|
|
||||||
sphinx_llms_txt._root_first_paragraph = ""
|
|
||||||
|
|
||||||
# Mock a Sphinx app with a disallowed builder
|
|
||||||
class MockBuilder:
|
|
||||||
name = "text" # Not in allowed list
|
|
||||||
|
|
||||||
class MockApp:
|
|
||||||
def __init__(self):
|
|
||||||
self.config_values = {}
|
|
||||||
self.connections = {}
|
|
||||||
self.builder = MockBuilder()
|
|
||||||
|
|
||||||
def add_config_value(self, name, default, rebuild):
|
|
||||||
self.config_values[name] = (default, rebuild)
|
|
||||||
|
|
||||||
def connect(self, event, handler):
|
|
||||||
self.connections[event] = handler
|
|
||||||
|
|
||||||
app = MockApp()
|
|
||||||
setup(app)
|
|
||||||
|
|
||||||
# Trigger builder-inited
|
|
||||||
builder_inited_handler = app.connections["builder-inited"]
|
|
||||||
builder_inited_handler(app)
|
|
||||||
|
|
||||||
# With disallowed builder, other events should NOT be connected
|
|
||||||
assert "doctree-resolved" not in app.connections
|
|
||||||
assert "build-finished" not in app.connections
|
|
||||||
|
|
||||||
|
|
||||||
def test_document_collector_initialization():
|
def test_document_collector_initialization():
|
||||||
"""Test initialization of DocumentCollector."""
|
"""Test initialization of DocumentCollector."""
|
||||||
collector = DocumentCollector()
|
collector = DocumentCollector()
|
||||||
@@ -370,664 +334,3 @@ def test_write_verbose_info_with_baseurl(tmp_path):
|
|||||||
|
|
||||||
assert "- [Home Page](https://example.org/index.html)" in content
|
assert "- [Home Page](https://example.org/index.html)" in content
|
||||||
assert "- [About Us](https://example.org/about.html)" in content
|
assert "- [About Us](https://example.org/about.html)" in content
|
||||||
|
|
||||||
|
|
||||||
def test_get_source_suffixes_with_dict():
|
|
||||||
"""Test _get_source_suffixes method with dict source_suffix."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with dict source_suffix
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
source_suffix = {".rst": None, ".md": None, ".txt": None}
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
|
|
||||||
suffixes = manager._get_source_suffixes()
|
|
||||||
assert set(suffixes) == {".rst", ".md", ".txt"}
|
|
||||||
|
|
||||||
|
|
||||||
def test_get_source_suffixes_with_list():
|
|
||||||
"""Test _get_source_suffixes method with list source_suffix."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with list source_suffix
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
source_suffix = [".rst", ".md"]
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
|
|
||||||
suffixes = manager._get_source_suffixes()
|
|
||||||
assert suffixes == [".rst", ".md"]
|
|
||||||
|
|
||||||
|
|
||||||
def test_get_source_suffixes_with_string():
|
|
||||||
"""Test _get_source_suffixes method with string source_suffix."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with string source_suffix
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
source_suffix = ".rst"
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
|
|
||||||
suffixes = manager._get_source_suffixes()
|
|
||||||
assert suffixes == [".rst"]
|
|
||||||
|
|
||||||
|
|
||||||
def test_get_source_suffixes_no_app():
|
|
||||||
"""Test _get_source_suffixes method with no app set."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
|
|
||||||
suffixes = manager._get_source_suffixes()
|
|
||||||
assert suffixes == [".rst"] # Default fallback
|
|
||||||
|
|
||||||
|
|
||||||
def test_html_sourcelink_suffix_default():
|
|
||||||
"""Test html_sourcelink_suffix defaults to .txt when no app is set."""
|
|
||||||
import tempfile
|
|
||||||
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_config(
|
|
||||||
{
|
|
||||||
"llms_txt_full_filename": "test.txt",
|
|
||||||
"llms_txt_exclude": [],
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create a temporary directory structure
|
|
||||||
with tempfile.TemporaryDirectory() as tmpdir:
|
|
||||||
outdir = f"{tmpdir}/build"
|
|
||||||
srcdir = f"{tmpdir}/source"
|
|
||||||
sources_dir = f"{outdir}/_sources"
|
|
||||||
|
|
||||||
# Create directories
|
|
||||||
import os
|
|
||||||
|
|
||||||
os.makedirs(sources_dir, exist_ok=True)
|
|
||||||
os.makedirs(srcdir, exist_ok=True)
|
|
||||||
|
|
||||||
# Create a test source file with default .txt suffix
|
|
||||||
test_file = f"{sources_dir}/index.rst.txt"
|
|
||||||
with open(test_file, "w") as f:
|
|
||||||
f.write("Test content")
|
|
||||||
|
|
||||||
# Mock env with minimal required attributes
|
|
||||||
class MockEnv:
|
|
||||||
all_docs = {"index": None}
|
|
||||||
titles = {
|
|
||||||
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
|
|
||||||
}
|
|
||||||
toctree_includes = {}
|
|
||||||
|
|
||||||
manager.set_env(MockEnv())
|
|
||||||
manager.set_master_doc("index")
|
|
||||||
|
|
||||||
# Test that it uses .txt as the default suffix
|
|
||||||
manager.combine_sources(outdir, srcdir)
|
|
||||||
|
|
||||||
# Verify the file was found and processed (check if output file exists)
|
|
||||||
output_file = f"{outdir}/test.txt"
|
|
||||||
assert os.path.exists(output_file)
|
|
||||||
|
|
||||||
|
|
||||||
def test_html_sourcelink_suffix_custom():
|
|
||||||
"""Test html_sourcelink_suffix uses custom value from Sphinx config."""
|
|
||||||
import tempfile
|
|
||||||
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with custom html_sourcelink_suffix
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
html_sourcelink_suffix = "source"
|
|
||||||
source_suffix = ".rst"
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
manager.set_config(
|
|
||||||
{
|
|
||||||
"llms_txt_full_filename": "test.txt",
|
|
||||||
"llms_txt_exclude": [],
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create a temporary directory structure
|
|
||||||
with tempfile.TemporaryDirectory() as tmpdir:
|
|
||||||
outdir = f"{tmpdir}/build"
|
|
||||||
srcdir = f"{tmpdir}/source"
|
|
||||||
sources_dir = f"{outdir}/_sources"
|
|
||||||
|
|
||||||
# Create directories
|
|
||||||
import os
|
|
||||||
|
|
||||||
os.makedirs(sources_dir, exist_ok=True)
|
|
||||||
os.makedirs(srcdir, exist_ok=True)
|
|
||||||
|
|
||||||
# Create a test source file with custom .source suffix
|
|
||||||
test_file = f"{sources_dir}/index.rst.source"
|
|
||||||
with open(test_file, "w") as f:
|
|
||||||
f.write("Test content")
|
|
||||||
|
|
||||||
# Mock env with minimal required attributes
|
|
||||||
class MockEnv:
|
|
||||||
all_docs = {"index": None}
|
|
||||||
titles = {
|
|
||||||
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
|
|
||||||
}
|
|
||||||
toctree_includes = {}
|
|
||||||
|
|
||||||
manager.set_env(MockEnv())
|
|
||||||
manager.set_master_doc("index")
|
|
||||||
|
|
||||||
# Test that it uses .source as the custom suffix
|
|
||||||
manager.combine_sources(outdir, srcdir)
|
|
||||||
|
|
||||||
# Verify the file was found and processed
|
|
||||||
output_file = f"{outdir}/test.txt"
|
|
||||||
assert os.path.exists(output_file)
|
|
||||||
|
|
||||||
|
|
||||||
def test_html_sourcelink_suffix_with_dot():
|
|
||||||
"""Test html_sourcelink_suffix adds dot if missing."""
|
|
||||||
import tempfile
|
|
||||||
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with html_sourcelink_suffix without leading dot
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
html_sourcelink_suffix = "src" # No leading dot
|
|
||||||
source_suffix = ".rst"
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
manager.set_config(
|
|
||||||
{
|
|
||||||
"llms_txt_full_filename": "test.txt",
|
|
||||||
"llms_txt_exclude": [],
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create a temporary directory structure
|
|
||||||
with tempfile.TemporaryDirectory() as tmpdir:
|
|
||||||
outdir = f"{tmpdir}/build"
|
|
||||||
srcdir = f"{tmpdir}/source"
|
|
||||||
sources_dir = f"{outdir}/_sources"
|
|
||||||
|
|
||||||
# Create directories
|
|
||||||
import os
|
|
||||||
|
|
||||||
os.makedirs(sources_dir, exist_ok=True)
|
|
||||||
os.makedirs(srcdir, exist_ok=True)
|
|
||||||
|
|
||||||
# Create a test source file with .src suffix (dot should be added automatically)
|
|
||||||
test_file = f"{sources_dir}/index.rst.src"
|
|
||||||
with open(test_file, "w") as f:
|
|
||||||
f.write("Test content")
|
|
||||||
|
|
||||||
# Mock env with minimal required attributes
|
|
||||||
class MockEnv:
|
|
||||||
all_docs = {"index": None}
|
|
||||||
titles = {
|
|
||||||
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
|
|
||||||
}
|
|
||||||
toctree_includes = {}
|
|
||||||
|
|
||||||
manager.set_env(MockEnv())
|
|
||||||
manager.set_master_doc("index")
|
|
||||||
|
|
||||||
# Test that it adds the dot and finds the file
|
|
||||||
manager.combine_sources(outdir, srcdir)
|
|
||||||
|
|
||||||
# Verify the file was found and processed
|
|
||||||
output_file = f"{outdir}/test.txt"
|
|
||||||
assert os.path.exists(output_file)
|
|
||||||
|
|
||||||
|
|
||||||
def test_mixed_source_file_formats():
|
|
||||||
"""Test handling of mixed source file formats (.rst, .md, .txt)."""
|
|
||||||
import tempfile
|
|
||||||
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with multiple source suffixes
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
html_sourcelink_suffix = ".txt"
|
|
||||||
source_suffix = {".rst": None, ".md": None, ".txt": None}
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
manager.set_config(
|
|
||||||
{
|
|
||||||
"llms_txt_full_filename": "test.txt",
|
|
||||||
"llms_txt_exclude": [],
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create a temporary directory structure
|
|
||||||
with tempfile.TemporaryDirectory() as tmpdir:
|
|
||||||
outdir = f"{tmpdir}/build"
|
|
||||||
srcdir = f"{tmpdir}/source"
|
|
||||||
sources_dir = f"{outdir}/_sources"
|
|
||||||
|
|
||||||
# Create directories
|
|
||||||
import os
|
|
||||||
|
|
||||||
os.makedirs(sources_dir, exist_ok=True)
|
|
||||||
os.makedirs(srcdir, exist_ok=True)
|
|
||||||
|
|
||||||
# Create test source files with different formats
|
|
||||||
files_to_create = [
|
|
||||||
f"{sources_dir}/page1.rst.txt",
|
|
||||||
f"{sources_dir}/page2.md.txt",
|
|
||||||
f"{sources_dir}/page3.txt.txt",
|
|
||||||
]
|
|
||||||
|
|
||||||
for test_file in files_to_create:
|
|
||||||
with open(test_file, "w") as f:
|
|
||||||
f.write(f"Content for {os.path.basename(test_file)}")
|
|
||||||
|
|
||||||
# Mock env with all documents
|
|
||||||
class MockEnv:
|
|
||||||
all_docs = {"page1": None, "page2": None, "page3": None}
|
|
||||||
titles = {
|
|
||||||
"page1": type("TitleNode", (), {"astext": lambda: "Page 1"})(),
|
|
||||||
"page2": type("TitleNode", (), {"astext": lambda: "Page 2"})(),
|
|
||||||
"page3": type("TitleNode", (), {"astext": lambda: "Page 3"})(),
|
|
||||||
}
|
|
||||||
toctree_includes = {}
|
|
||||||
|
|
||||||
manager.set_env(MockEnv())
|
|
||||||
manager.set_master_doc("page1")
|
|
||||||
|
|
||||||
# Test that all file formats are found and processed
|
|
||||||
manager.combine_sources(outdir, srcdir)
|
|
||||||
|
|
||||||
# Verify the output file was created and contains content from all formats
|
|
||||||
output_file = f"{outdir}/test.txt"
|
|
||||||
assert os.path.exists(output_file)
|
|
||||||
|
|
||||||
with open(output_file, "r") as f:
|
|
||||||
content = f.read()
|
|
||||||
|
|
||||||
# Should contain content from all three files
|
|
||||||
assert "Content for page1.rst.txt" in content
|
|
||||||
assert "Content for page2.md.txt" in content
|
|
||||||
assert "Content for page3.txt.txt" in content
|
|
||||||
|
|
||||||
|
|
||||||
def test_source_suffix_detection_priority():
|
|
||||||
"""Test source suffix detection tries formats in correct order for docnames."""
|
|
||||||
import tempfile
|
|
||||||
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Mock Sphinx app with ordered source suffixes
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
html_sourcelink_suffix = ".txt"
|
|
||||||
source_suffix = [".rst", ".md"] # rst has priority over md
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.set_app(MockApp())
|
|
||||||
manager.set_config(
|
|
||||||
{
|
|
||||||
"llms_txt_full_filename": "test.txt",
|
|
||||||
"llms_txt_exclude": [],
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create a temporary directory structure
|
|
||||||
with tempfile.TemporaryDirectory() as tmpdir:
|
|
||||||
outdir = f"{tmpdir}/build"
|
|
||||||
srcdir = f"{tmpdir}/source"
|
|
||||||
sources_dir = f"{outdir}/_sources"
|
|
||||||
|
|
||||||
# Create directories
|
|
||||||
import os
|
|
||||||
|
|
||||||
os.makedirs(sources_dir, exist_ok=True)
|
|
||||||
os.makedirs(srcdir, exist_ok=True)
|
|
||||||
|
|
||||||
# Create both .rst and .md versions of the same document
|
|
||||||
# Only create files for the specific docname "index"
|
|
||||||
rst_file = f"{sources_dir}/index.rst.txt"
|
|
||||||
md_file = f"{sources_dir}/index.md.txt"
|
|
||||||
|
|
||||||
with open(rst_file, "w") as f:
|
|
||||||
f.write("RST content for index")
|
|
||||||
|
|
||||||
with open(md_file, "w") as f:
|
|
||||||
f.write("Markdown content for index")
|
|
||||||
|
|
||||||
# Mock env with only the index document
|
|
||||||
class MockEnv:
|
|
||||||
all_docs = {"index": None}
|
|
||||||
titles = {
|
|
||||||
"index": type("TitleNode", (), {"astext": lambda: "Index Page"})()
|
|
||||||
}
|
|
||||||
toctree_includes = {"index": []}
|
|
||||||
|
|
||||||
manager.set_env(MockEnv())
|
|
||||||
manager.set_master_doc("index")
|
|
||||||
|
|
||||||
# Test the priority behavior
|
|
||||||
manager.combine_sources(outdir, srcdir)
|
|
||||||
|
|
||||||
# Check that output file was created
|
|
||||||
output_file = f"{outdir}/test.txt"
|
|
||||||
assert os.path.exists(output_file)
|
|
||||||
|
|
||||||
with open(output_file, "r") as f:
|
|
||||||
content = f.read()
|
|
||||||
|
|
||||||
# The system should prefer RST over MD for the "index" docname
|
|
||||||
# But since both files exist and the second phase adds remaining files,
|
|
||||||
# both will be included. The test verifies that RST appears first
|
|
||||||
# (indicating it was found first in the priority order)
|
|
||||||
assert "RST content for index" in content
|
|
||||||
|
|
||||||
# Find positions to verify order
|
|
||||||
rst_pos = content.find("RST content for index")
|
|
||||||
md_pos = content.find("Markdown content for index")
|
|
||||||
|
|
||||||
# RST should come before MD (due to priority in toctree processing)
|
|
||||||
assert rst_pos < md_pos, "RST content should appear before MD content"
|
|
||||||
|
|
||||||
|
|
||||||
def test_summary_default_uses_first_paragraph():
|
|
||||||
"""
|
|
||||||
Test that summary defaults to first paragraph of root document when not configured.
|
|
||||||
"""
|
|
||||||
from docutils import nodes
|
|
||||||
from docutils.frontend import OptionParser
|
|
||||||
from docutils.parsers.rst import Parser
|
|
||||||
from docutils.utils import new_document
|
|
||||||
|
|
||||||
from sphinx_llms_txt import build_finished, doctree_resolved
|
|
||||||
|
|
||||||
# Create a proper document with settings
|
|
||||||
settings = OptionParser(components=(Parser,)).get_default_values()
|
|
||||||
doctree = new_document("<rst-doc>", settings)
|
|
||||||
|
|
||||||
title = nodes.title(text="Test Title")
|
|
||||||
paragraph = nodes.paragraph(
|
|
||||||
text="This is the first paragraph that should be used as summary."
|
|
||||||
)
|
|
||||||
doctree.append(title)
|
|
||||||
doctree.append(paragraph)
|
|
||||||
|
|
||||||
# Mock Sphinx app
|
|
||||||
class MockApp:
|
|
||||||
class Config:
|
|
||||||
master_doc = "index"
|
|
||||||
llms_txt_summary = None # Not configured
|
|
||||||
llms_txt_file = True
|
|
||||||
llms_txt_filename = "llms.txt"
|
|
||||||
llms_txt_title = None
|
|
||||||
llms_txt_full_file = True
|
|
||||||
llms_txt_full_filename = "llms-full.txt"
|
|
||||||
llms_txt_full_max_size = None
|
|
||||||
llms_txt_full_size_policy = "warn_skip"
|
|
||||||
llms_txt_directives = []
|
|
||||||
llms_txt_exclude = []
|
|
||||||
llms_txt_code_files = []
|
|
||||||
llms_txt_code_base_path = None
|
|
||||||
html_baseurl = ""
|
|
||||||
|
|
||||||
config = Config()
|
|
||||||
outdir = "/tmp/build"
|
|
||||||
srcdir = "/tmp/source"
|
|
||||||
|
|
||||||
class Env:
|
|
||||||
titles = {
|
|
||||||
"index": type("TitleNode", (), {"astext": lambda self: "Test Title"})()
|
|
||||||
}
|
|
||||||
|
|
||||||
env = Env()
|
|
||||||
|
|
||||||
app = MockApp()
|
|
||||||
|
|
||||||
# Reset the global state
|
|
||||||
import sphinx_llms_txt
|
|
||||||
|
|
||||||
sphinx_llms_txt._root_first_paragraph = ""
|
|
||||||
|
|
||||||
# Call doctree_resolved to extract the first paragraph
|
|
||||||
doctree_resolved(app, doctree, "index")
|
|
||||||
|
|
||||||
# Verify the first paragraph was extracted
|
|
||||||
assert (
|
|
||||||
sphinx_llms_txt._root_first_paragraph
|
|
||||||
== "This is the first paragraph that should be used as summary."
|
|
||||||
)
|
|
||||||
|
|
||||||
# Mock the manager methods to avoid actual file operations
|
|
||||||
original_combine_sources = sphinx_llms_txt._manager.combine_sources
|
|
||||||
sphinx_llms_txt._manager.combine_sources = lambda outdir, srcdir: None
|
|
||||||
|
|
||||||
# Call build_finished and verify the summary is set correctly
|
|
||||||
build_finished(app, None)
|
|
||||||
|
|
||||||
# Check that the summary was properly configured
|
|
||||||
assert (
|
|
||||||
sphinx_llms_txt._manager.config["llms_txt_summary"]
|
|
||||||
== "This is the first paragraph that should be used as summary."
|
|
||||||
)
|
|
||||||
|
|
||||||
# Restore original method
|
|
||||||
sphinx_llms_txt._manager.combine_sources = original_combine_sources
|
|
||||||
|
|
||||||
|
|
||||||
def test_code_files_include_exclude_patterns(tmp_path):
|
|
||||||
"""Test the +/- pattern syntax for llms_txt_code_files configuration."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Create test directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
|
|
||||||
docs_dir = src_dir / "docs"
|
|
||||||
docs_dir.mkdir()
|
|
||||||
|
|
||||||
cache_dir = docs_dir / "__pycache__"
|
|
||||||
cache_dir.mkdir()
|
|
||||||
|
|
||||||
# Create test files
|
|
||||||
(docs_dir / "example.rst").write_text("Example RST content")
|
|
||||||
(docs_dir / "guide.rst").write_text("Guide RST content")
|
|
||||||
(docs_dir / "backup.bak").write_text("Backup file content")
|
|
||||||
(cache_dir / "compiled.pyc").write_text("Compiled Python")
|
|
||||||
|
|
||||||
# Create manager and set source directory
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Test configuration with include/exclude patterns
|
|
||||||
config = {
|
|
||||||
"llms_txt_code_files": [
|
|
||||||
"+:docs/**/*.rst", # Include all RST files in docs
|
|
||||||
"-:docs/**/__pycache__/**", # Exclude pycache files
|
|
||||||
"-:docs/**/*.bak", # Exclude backup files
|
|
||||||
]
|
|
||||||
}
|
|
||||||
manager.set_config(config)
|
|
||||||
|
|
||||||
# Process code files
|
|
||||||
code_parts, _ = manager._process_code_files()
|
|
||||||
|
|
||||||
# Verify we have the expected number of files
|
|
||||||
assert len(code_parts) == 2, f"Expected 2 files, got {len(code_parts)}"
|
|
||||||
|
|
||||||
# Extract file titles from code blocks
|
|
||||||
titles = []
|
|
||||||
for part in code_parts:
|
|
||||||
lines = part.strip().split("\n")
|
|
||||||
if lines:
|
|
||||||
titles.append(lines[0])
|
|
||||||
|
|
||||||
# Verify expected files are included
|
|
||||||
assert "docs/example.rst" in titles
|
|
||||||
assert "docs/guide.rst" in titles
|
|
||||||
|
|
||||||
# Verify excluded files are not present
|
|
||||||
content = "\n".join(code_parts)
|
|
||||||
assert "backup.bak" not in content
|
|
||||||
assert "__pycache__" not in content
|
|
||||||
assert "compiled.pyc" not in content
|
|
||||||
|
|
||||||
|
|
||||||
def test_code_files_exclude_only_patterns(tmp_path):
|
|
||||||
"""Test that exclude-only patterns result in no files being included."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Create test directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
|
|
||||||
docs_dir = src_dir / "docs"
|
|
||||||
docs_dir.mkdir()
|
|
||||||
|
|
||||||
# Create test files
|
|
||||||
(docs_dir / "example.rst").write_text("Example RST content")
|
|
||||||
|
|
||||||
# Create manager and set source directory
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Test configuration with only exclude patterns
|
|
||||||
config = {
|
|
||||||
"llms_txt_code_files": [
|
|
||||||
"-:docs/**/*.rst", # Only exclude pattern, no includes
|
|
||||||
]
|
|
||||||
}
|
|
||||||
manager.set_config(config)
|
|
||||||
|
|
||||||
# Process code files
|
|
||||||
code_parts, _ = manager._process_code_files()
|
|
||||||
|
|
||||||
# Should have no files with exclude-only patterns
|
|
||||||
assert len(code_parts) == 0, "Should have no files with exclude-only patterns"
|
|
||||||
|
|
||||||
|
|
||||||
def test_code_files_no_prefix_patterns(tmp_path):
|
|
||||||
"""Test that patterns without prefix are ignored."""
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Create test directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
|
|
||||||
docs_dir = src_dir / "docs"
|
|
||||||
docs_dir.mkdir()
|
|
||||||
|
|
||||||
# Create test files
|
|
||||||
(docs_dir / "example.rst").write_text("Example RST content")
|
|
||||||
(docs_dir / "backup.bak").write_text("Backup file content")
|
|
||||||
|
|
||||||
# Create manager and set source directory
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Test configuration with no prefix (should be ignored)
|
|
||||||
config = {
|
|
||||||
"llms_txt_code_files": [
|
|
||||||
"docs/**/*.rst", # No prefix = ignored
|
|
||||||
"+:docs/**/*.rst", # Include RST files
|
|
||||||
"-:docs/**/*.bak", # Exclude backup files
|
|
||||||
]
|
|
||||||
}
|
|
||||||
manager.set_config(config)
|
|
||||||
|
|
||||||
# Process code files
|
|
||||||
code_parts, _ = manager._process_code_files()
|
|
||||||
|
|
||||||
# Should include RST files (from +: pattern) and exclude BAK files (from -: pattern)
|
|
||||||
assert len(code_parts) == 1, "Should include RST files and exclude BAK files"
|
|
||||||
|
|
||||||
content = "\n".join(code_parts)
|
|
||||||
assert "Example RST content" in content
|
|
||||||
assert "backup.bak" not in content
|
|
||||||
|
|
||||||
|
|
||||||
def test_code_files_ignored_patterns(tmp_path, caplog):
|
|
||||||
"""Test that patterns without +: or -: prefix log a warning and are ignored."""
|
|
||||||
from unittest.mock import patch
|
|
||||||
|
|
||||||
from sphinx_llms_txt.manager import LLMSFullManager
|
|
||||||
|
|
||||||
# Create test directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
|
|
||||||
docs_dir = src_dir / "docs"
|
|
||||||
docs_dir.mkdir()
|
|
||||||
|
|
||||||
# Create test files
|
|
||||||
(docs_dir / "example.rst").write_text("Example RST content")
|
|
||||||
|
|
||||||
# Use a mock to capture the warning message directly
|
|
||||||
captured_warnings = []
|
|
||||||
|
|
||||||
def capture_warning(message, *args, **kwargs):
|
|
||||||
captured_warnings.append(message)
|
|
||||||
|
|
||||||
# Patch the logger to capture warnings
|
|
||||||
with patch("sphinx_llms_txt.manager.logger.warning", side_effect=capture_warning):
|
|
||||||
# Create manager and set source directory
|
|
||||||
manager = LLMSFullManager()
|
|
||||||
manager.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Test configuration with only no-prefix patterns (should result in no files)
|
|
||||||
config = {
|
|
||||||
"llms_txt_code_files": [
|
|
||||||
"docs/**/*.rst", # No prefix = ignored with warning
|
|
||||||
]
|
|
||||||
}
|
|
||||||
manager.set_config(config)
|
|
||||||
|
|
||||||
# Process code files
|
|
||||||
code_parts, _ = manager._process_code_files()
|
|
||||||
|
|
||||||
# Should have no files since the pattern without prefix is ignored
|
|
||||||
assert (
|
|
||||||
len(code_parts) == 0
|
|
||||||
), "Should have no files when only using patterns without prefix"
|
|
||||||
|
|
||||||
# Check that a warning was logged
|
|
||||||
assert (
|
|
||||||
len(captured_warnings) == 1
|
|
||||||
), f"Expected 1 warning, got {len(captured_warnings)}"
|
|
||||||
assert (
|
|
||||||
"Code file pattern 'docs/**/*.rst' ignored." in captured_warnings[0]
|
|
||||||
), f"Warning message should contain expected text. Got: {captured_warnings[0]}"
|
|
||||||
|
|||||||
@@ -102,7 +102,7 @@ def test_process_path_directives_with_html_baseurl(tmp_path):
|
|||||||
|
|
||||||
|
|
||||||
def test_process_path_directives_absolute_urls(tmp_path):
|
def test_process_path_directives_absolute_urls(tmp_path):
|
||||||
"""Test that absolute URLs are not modified but absolute paths get base URL."""
|
"""Test that absolute URLs are not modified."""
|
||||||
# Create a processor
|
# Create a processor
|
||||||
config = {
|
config = {
|
||||||
"llms_txt_directives": [],
|
"llms_txt_directives": [],
|
||||||
@@ -127,17 +127,10 @@ def test_process_path_directives_absolute_urls(tmp_path):
|
|||||||
with open(source_file, "w", encoding="utf-8") as f:
|
with open(source_file, "w", encoding="utf-8") as f:
|
||||||
f.write(source_content)
|
f.write(source_content)
|
||||||
|
|
||||||
# Process the directives
|
# Process the directives (should remain unchanged)
|
||||||
processed_content = processor._process_path_directives(source_content, source_file)
|
processed_content = processor._process_path_directives(source_content, source_file)
|
||||||
|
|
||||||
# Expected: URLs and data URIs unchanged, absolute paths get base URL
|
assert processed_content == source_content
|
||||||
expected_content = (
|
|
||||||
".. image:: https://othersite.com/images/test.png\n"
|
|
||||||
".. image:: https://example.com/docs/absolute/path/image.png\n"
|
|
||||||
".. image:: data:image/png;base64,iVBORw0KG...\n"
|
|
||||||
)
|
|
||||||
|
|
||||||
assert processed_content == expected_content
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_path_directives_custom_directives(tmp_path):
|
def test_process_path_directives_custom_directives(tmp_path):
|
||||||
@@ -258,149 +251,3 @@ def test_process_content_end_to_end(tmp_path):
|
|||||||
)
|
)
|
||||||
|
|
||||||
assert processed_content == expected_content
|
assert processed_content == expected_content
|
||||||
|
|
||||||
|
|
||||||
def test_process_path_directives_images_directory(tmp_path):
|
|
||||||
"""Test that _images directory paths are handled correctly."""
|
|
||||||
# Create a processor with base URL
|
|
||||||
config = {
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
"html_baseurl": "https://example.com/docs",
|
|
||||||
}
|
|
||||||
processor = DocumentProcessor(config)
|
|
||||||
|
|
||||||
# Create source directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
processor.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Create _sources directory to mimic Sphinx output
|
|
||||||
build_dir = tmp_path / "build"
|
|
||||||
build_dir.mkdir()
|
|
||||||
sources_dir = build_dir / "_sources"
|
|
||||||
sources_dir.mkdir()
|
|
||||||
|
|
||||||
# Create a source file with various _images directory paths
|
|
||||||
source_content = (
|
|
||||||
"Some content.\n"
|
|
||||||
".. image:: _images/test.png\n" # Relative _images should become /_images
|
|
||||||
".. image:: /_images/absolute.png\n" # Absolute _images should get base URL
|
|
||||||
".. figure:: _images/figure.png\n" # Test with figure directive too
|
|
||||||
" :alt: A test figure\n"
|
|
||||||
".. image:: images/normal.png\n" # Normal relative path should be unchanged
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create source file in sources directory to simulate Sphinx build output
|
|
||||||
source_file = sources_dir / "page.txt"
|
|
||||||
with open(source_file, "w", encoding="utf-8") as f:
|
|
||||||
f.write(source_content)
|
|
||||||
|
|
||||||
# Process the directives
|
|
||||||
processed_content = processor._process_path_directives(source_content, source_file)
|
|
||||||
|
|
||||||
# Expected: _images paths should be converted and get base URL
|
|
||||||
expected_content = (
|
|
||||||
"Some content.\n"
|
|
||||||
".. image:: https://example.com/docs/_images/test.png\n"
|
|
||||||
".. image:: https://example.com/docs/_images/absolute.png\n"
|
|
||||||
".. figure:: https://example.com/docs/_images/figure.png\n"
|
|
||||||
" :alt: A test figure\n"
|
|
||||||
".. image:: https://example.com/docs/images/normal.png\n"
|
|
||||||
)
|
|
||||||
|
|
||||||
assert processed_content == expected_content
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_path_directives_images_directory_no_baseurl(tmp_path):
|
|
||||||
"""
|
|
||||||
Test that _images directory paths work correctly without base URL.
|
|
||||||
Only converts when image exists.
|
|
||||||
"""
|
|
||||||
# Create a processor without base URL
|
|
||||||
config = {
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
"html_baseurl": "",
|
|
||||||
}
|
|
||||||
processor = DocumentProcessor(config)
|
|
||||||
|
|
||||||
# Create source directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
processor.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Create _sources directory to mimic Sphinx output
|
|
||||||
build_dir = tmp_path / "build"
|
|
||||||
build_dir.mkdir()
|
|
||||||
sources_dir = build_dir / "_sources"
|
|
||||||
sources_dir.mkdir()
|
|
||||||
|
|
||||||
# Create _images directory and one test image
|
|
||||||
images_dir = build_dir / "_images"
|
|
||||||
images_dir.mkdir()
|
|
||||||
(images_dir / "test.png").write_text("fake image content")
|
|
||||||
# Note: absolute.png is not created, so it won't be converted
|
|
||||||
|
|
||||||
# Create a source file with _images directory paths
|
|
||||||
source_content = (
|
|
||||||
".. image:: _images/test.png\n" # Should become /_images (image exists)
|
|
||||||
".. image:: /_images/absolute.png\n" # Should stay unchanged (absolute path)
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create source file in sources directory to simulate Sphinx build output
|
|
||||||
source_file = sources_dir / "page.txt"
|
|
||||||
with open(source_file, "w", encoding="utf-8") as f:
|
|
||||||
f.write(source_content)
|
|
||||||
|
|
||||||
# Process the directives
|
|
||||||
processed_content = processor._process_path_directives(source_content, source_file)
|
|
||||||
|
|
||||||
# Expected: only test.png gets converted because it exists in _images
|
|
||||||
expected_content = (
|
|
||||||
".. image:: /_images/test.png\n" # Converted because image exists
|
|
||||||
".. image:: /_images/absolute.png\n" # Absolute path unchanged
|
|
||||||
)
|
|
||||||
|
|
||||||
assert processed_content == expected_content
|
|
||||||
|
|
||||||
|
|
||||||
def test_process_path_directives_all_absolute_paths_get_baseurl(tmp_path):
|
|
||||||
"""Test that all absolute paths (starting with /) get base URL prepended."""
|
|
||||||
# Create a processor with base URL
|
|
||||||
config = {
|
|
||||||
"llms_txt_directives": [],
|
|
||||||
"html_baseurl": "https://mysite.com/docs/",
|
|
||||||
}
|
|
||||||
processor = DocumentProcessor(config)
|
|
||||||
|
|
||||||
# Create source directory structure
|
|
||||||
src_dir = tmp_path / "src"
|
|
||||||
src_dir.mkdir()
|
|
||||||
processor.srcdir = str(src_dir)
|
|
||||||
|
|
||||||
# Create a source file with various absolute paths
|
|
||||||
source_content = (
|
|
||||||
".. image:: /static/images/logo.png\n"
|
|
||||||
".. figure:: /assets/diagrams/flow.svg\n"
|
|
||||||
".. image:: /media/photos/team.jpg\n"
|
|
||||||
" :alt: Team photo\n"
|
|
||||||
".. image:: relative/path.png\n" # This should still get normal processing
|
|
||||||
)
|
|
||||||
|
|
||||||
# Create source file
|
|
||||||
source_file = src_dir / "page.txt"
|
|
||||||
with open(source_file, "w", encoding="utf-8") as f:
|
|
||||||
f.write(source_content)
|
|
||||||
|
|
||||||
# Process the directives
|
|
||||||
processed_content = processor._process_path_directives(source_content, source_file)
|
|
||||||
|
|
||||||
# Expected: All absolute paths get base URL prepended
|
|
||||||
expected_content = (
|
|
||||||
".. image:: https://mysite.com/docs/static/images/logo.png\n"
|
|
||||||
".. figure:: https://mysite.com/docs/assets/diagrams/flow.svg\n"
|
|
||||||
".. image:: https://mysite.com/docs/media/photos/team.jpg\n"
|
|
||||||
" :alt: Team photo\n"
|
|
||||||
".. image:: https://mysite.com/docs/relative/path.png\n"
|
|
||||||
)
|
|
||||||
|
|
||||||
assert processed_content == expected_content
|
|
||||||
|
|||||||
Reference in New Issue
Block a user