Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
46c2dec254 | ||
|
|
b56d93d265 | ||
|
|
5db5395889 | ||
|
|
04b0657dc5 | ||
|
|
935a964c7e | ||
|
|
51f6c71de3 | ||
|
|
70defd3996 | ||
|
|
9ae05c6c13 | ||
|
|
5581979cac | ||
|
|
f70f1a26ec | ||
|
|
ed50138ae4 | ||
|
|
8f4d2c07c6 |
@@ -1,6 +1,28 @@
|
|||||||
Changelog
|
Changelog
|
||||||
=========
|
=========
|
||||||
|
|
||||||
|
0.3.0
|
||||||
|
-----
|
||||||
|
|
||||||
|
- Use first paragraph as default for ``llms_txt_summary``
|
||||||
|
`#22 <https://github.com/jdillard/sphinx-llms-txt/pull/22>`_
|
||||||
|
|
||||||
|
0.2.4
|
||||||
|
-----
|
||||||
|
|
||||||
|
- Support source file suffix detection
|
||||||
|
`#21 <https://github.com/jdillard/sphinx-llms-txt/pull/21>`_
|
||||||
|
|
||||||
|
0.2.3
|
||||||
|
-----
|
||||||
|
|
||||||
|
- Remove ``get_and_resolve_toctree`` method
|
||||||
|
`#19 <https://github.com/jdillard/sphinx-llms-txt/pull/19>`_
|
||||||
|
- Simplify ``_sources`` lookup
|
||||||
|
`#18 <https://github.com/jdillard/sphinx-llms-txt/pull/18>`_
|
||||||
|
- Add sphinx docs
|
||||||
|
`#16 <https://github.com/jdillard/sphinx-llms-txt/pull/16>`_
|
||||||
|
|
||||||
0.2.2
|
0.2.2
|
||||||
-----
|
-----
|
||||||
|
|
||||||
|
|||||||
@@ -1,91 +1,14 @@
|
|||||||
# Sphinx llms.txt generator
|
# Sphinx llms.txt generator
|
||||||
|
|
||||||
A Sphinx extension that generates a summary `llms.txt` file, written in Markdown, and a single combined documentation `llms-full.txt` file, written in reStructuredText.
|
A Sphinx extension that generates a summary `llms.txt` file and a single combined documentation `llms-full.txt` file.
|
||||||
|
|
||||||
[](https://pypi.python.org/pypi/sphinx-llms-txt)
|
[](https://pypi.python.org/pypi/sphinx-llms-txt)
|
||||||
[](https://pepy.tech/project/sphinx-llms-txt)
|
[](https://pepy.tech/project/sphinx-llms-txt)
|
||||||
|
[](#)
|
||||||
|
|
||||||
## Installation
|
## Documentation
|
||||||
|
|
||||||
```bash
|
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
|
||||||
pip install sphinx-llms-txt
|
|
||||||
```
|
|
||||||
|
|
||||||
## Usage
|
|
||||||
|
|
||||||
1. Add the extension to your Sphinx configuration (`conf.py`):
|
|
||||||
|
|
||||||
```python
|
|
||||||
extensions = [
|
|
||||||
'sphinx_llms_txt',
|
|
||||||
]
|
|
||||||
```
|
|
||||||
|
|
||||||
## Configuration Options
|
|
||||||
|
|
||||||
### `llms_txt_full_file`
|
|
||||||
|
|
||||||
- **Type**: boolean
|
|
||||||
- **Default**: `'True'`
|
|
||||||
- **Description**: Whether to write the single output file
|
|
||||||
|
|
||||||
### `llms_txt_full_filename`
|
|
||||||
|
|
||||||
- **Type**: string
|
|
||||||
- **Default**: `'llms-full.txt'`
|
|
||||||
- **Description**: Name of the single output file
|
|
||||||
|
|
||||||
### `llms_txt_full_max_size`
|
|
||||||
|
|
||||||
- **Type**: integer or `None`
|
|
||||||
- **Default**: `None` (no limit)
|
|
||||||
- **Description**: Sets a maximum line count for `llms_txt_full_filename`.
|
|
||||||
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
|
||||||
|
|
||||||
### `llms_txt_file`
|
|
||||||
|
|
||||||
- **Type**: boolean
|
|
||||||
- **Default**: `True`
|
|
||||||
- **Description**: Whether to write the summary information file
|
|
||||||
|
|
||||||
### `llms_txt_filename`
|
|
||||||
|
|
||||||
- **Type**: string
|
|
||||||
- **Default**: `llms.txt`
|
|
||||||
- **Description**: Name of the summary information file
|
|
||||||
|
|
||||||
### `llms_txt_directives`
|
|
||||||
|
|
||||||
- **Type**: list of strings
|
|
||||||
- **Default**: `[]`
|
|
||||||
- **Description**: List of custom directive names to process for path resolution.
|
|
||||||
|
|
||||||
### `llms_txt_title`
|
|
||||||
|
|
||||||
- **Type**: string or `None`
|
|
||||||
- **Default**: `None`
|
|
||||||
- **Description**: Overrides the Sphinx project name as the heading in `llms.txt`.
|
|
||||||
|
|
||||||
### `llms_txt_summary`
|
|
||||||
|
|
||||||
- **Type**: string or `None`
|
|
||||||
- **Default**: `None`
|
|
||||||
- **Description**: Optional, but recommended, summary description for `llms.txt`.
|
|
||||||
|
|
||||||
### `llms_txt_exclude`
|
|
||||||
|
|
||||||
- **Type**: list of strings
|
|
||||||
- **Default**: `[]`
|
|
||||||
- **Description**: A list of pages to ignore (e.g., `["page1", "page_with_*"]`).
|
|
||||||
|
|
||||||
## Features
|
|
||||||
|
|
||||||
- Creates `llms.txt` and `llms-full.txt`
|
|
||||||
- Automatically add content from `include` directives
|
|
||||||
- Resolves relative paths in directives like `image` and `figure` to use full paths
|
|
||||||
- Ability to add list of custom directives with `llms_txt_directives`
|
|
||||||
- Optionally, prepend a base URL using Sphinx's `html_baseurl`
|
|
||||||
- Ability to exclude pages
|
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,170 @@
|
|||||||
|
Advanced Configuration
|
||||||
|
======================
|
||||||
|
|
||||||
|
This page covers advanced configuration options for the sphinx-llms-txt extension.
|
||||||
|
|
||||||
|
.. _customizing_llms_files:
|
||||||
|
|
||||||
|
Customizing the LLMs Files
|
||||||
|
^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
By default, the extension generates two files:
|
||||||
|
|
||||||
|
1. ``llms.txt`` - A summary file in Markdown format
|
||||||
|
2. ``llms-full.txt`` - A complete documentation file in reStructuredText format
|
||||||
|
|
||||||
|
You can customize these files in several ways:
|
||||||
|
|
||||||
|
.. _changing_filenames:
|
||||||
|
|
||||||
|
Changing Filenames
|
||||||
|
~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
You can change the default filenames by setting these values in your ``conf.py``:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_filename = "custom-summary.txt"
|
||||||
|
llms_txt_full_filename = "custom-docs.txt"
|
||||||
|
|
||||||
|
.. _disabling_file_generation:
|
||||||
|
|
||||||
|
Disabling File Generation
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
If you only want one of the files, you can disable generation of the other:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
# Disable summary file
|
||||||
|
llms_txt_file = False
|
||||||
|
|
||||||
|
# Disable full documentation file
|
||||||
|
llms_txt_full_file = False
|
||||||
|
|
||||||
|
.. _custom_summary:
|
||||||
|
|
||||||
|
Adding a Custom Summary
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
The summary file can include a custom description of your project:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_summary = """
|
||||||
|
This documentation explains how to use MyProject to build amazing
|
||||||
|
applications. The project provides a comprehensive API for handling
|
||||||
|
data processing and visualization.
|
||||||
|
"""
|
||||||
|
|
||||||
|
.. note:: The summary can span multiple lines and will be properly formatted in the output file.
|
||||||
|
|
||||||
|
.. _custom_title:
|
||||||
|
|
||||||
|
Custom Title
|
||||||
|
~~~~~~~~~~~~
|
||||||
|
|
||||||
|
By default, the project name from Sphinx is used as the title in ``llms.txt``. You can override this:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_title = "My Custom Project Documentation"
|
||||||
|
|
||||||
|
.. _handling_large_documentation:
|
||||||
|
|
||||||
|
Handling Large Documentation
|
||||||
|
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
|
||||||
|
You can set a maximum line count:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
|
||||||
|
|
||||||
|
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
|
||||||
|
|
||||||
|
.. tip:: Use :ref:`excluding_content` to remove less relevant pages.
|
||||||
|
|
||||||
|
.. _custom_directive_handling:
|
||||||
|
|
||||||
|
Custom Directive Handling
|
||||||
|
^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
.. _path_resolution:
|
||||||
|
|
||||||
|
Path Resolution
|
||||||
|
~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
The extension resolves paths in the common directives ``[ 'image', 'figure']`` by default.
|
||||||
|
You can add custom directives to this list:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_directives = [
|
||||||
|
"my-custom-image-directive",
|
||||||
|
"another-directive-with-paths",
|
||||||
|
]
|
||||||
|
|
||||||
|
This ensures that paths in your custom directives are properly resolved in the generated files.
|
||||||
|
|
||||||
|
.. _excluding_content:
|
||||||
|
|
||||||
|
Excluding Content
|
||||||
|
^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
You can exclude specific pages from being included in the generated files:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
llms_txt_exclude = [
|
||||||
|
"search", # Exclude the search page
|
||||||
|
"genindex", # Exclude the index page
|
||||||
|
"private_*", # Exclude all pages starting with 'private_'
|
||||||
|
]
|
||||||
|
|
||||||
|
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
|
||||||
|
|
||||||
|
.. _using_html_baseurl:
|
||||||
|
|
||||||
|
Using HTML Base URL
|
||||||
|
^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
If you want to include absolute URLs for resources in your documentation, you can use Sphinx's built-in ``html_baseurl`` configuration:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
html_baseurl = "https://example.com/docs/"
|
||||||
|
|
||||||
|
When this option is set, all resolved paths in directives will be prefixed with this URL, creating absolute paths in the generated files.
|
||||||
|
|
||||||
|
.. _integration_examples:
|
||||||
|
|
||||||
|
Integration Examples
|
||||||
|
^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
Complete Configuration Example
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
Here's a complete example showing multiple :doc:`configuration-values`:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
# File names and generation options
|
||||||
|
llms_txt_filename = "ai-summary.txt"
|
||||||
|
llms_txt_full_filename = "ai-full-docs.txt"
|
||||||
|
llms_txt_full_max_size = 50000
|
||||||
|
|
||||||
|
# Content customization
|
||||||
|
llms_txt_title = "Project Documentation for AI Assistants"
|
||||||
|
llms_txt_summary = """
|
||||||
|
This is a comprehensive documentation set for our project.
|
||||||
|
It includes API references, usage examples, and tutorials.
|
||||||
|
"""
|
||||||
|
|
||||||
|
# Path handling
|
||||||
|
html_baseurl = "https://docs.example.com/"
|
||||||
|
llms_txt_directives = ["custom-image", "custom-include"]
|
||||||
|
|
||||||
|
# Content filtering
|
||||||
|
llms_txt_exclude = ["search", "genindex", "404", "private_*"]
|
||||||
+1
-1
@@ -12,7 +12,7 @@ import subprocess
|
|||||||
|
|
||||||
# -- Project information -----------------------------------------------------
|
# -- Project information -----------------------------------------------------
|
||||||
|
|
||||||
project = "Sphinx llms.txt Generator"
|
project = "sphinx-llms-txt"
|
||||||
copyright = "Jared Dillard"
|
copyright = "Jared Dillard"
|
||||||
author = "Jared Dillard"
|
author = "Jared Dillard"
|
||||||
llms_txt_summary = """
|
llms_txt_summary = """
|
||||||
|
|||||||
@@ -5,7 +5,8 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: boolean
|
- **Type**: boolean
|
||||||
- **Default**: ``True``
|
- **Default**: ``True``
|
||||||
- **Description**: Whether to write the single output file
|
- **Description**: Whether to write the single output file.
|
||||||
|
See :ref:`disabling_file_generation`.
|
||||||
|
|
||||||
.. versionadded:: 0.1.0
|
.. versionadded:: 0.1.0
|
||||||
|
|
||||||
@@ -13,7 +14,8 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: string
|
- **Type**: string
|
||||||
- **Default**: ``'llms-full.txt'``
|
- **Default**: ``'llms-full.txt'``
|
||||||
- **Description**: Name of the single output file
|
- **Description**: Name of the single output file.
|
||||||
|
See :ref:`changing_filenames`.
|
||||||
|
|
||||||
.. versionadded:: 0.1.0
|
.. versionadded:: 0.1.0
|
||||||
|
|
||||||
@@ -23,6 +25,7 @@ Project Configuration Values
|
|||||||
- **Default**: ``None`` (no limit)
|
- **Default**: ``None`` (no limit)
|
||||||
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
||||||
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
||||||
|
See :ref:`handling_large_documentation`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
@@ -30,7 +33,8 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: boolean
|
- **Type**: boolean
|
||||||
- **Default**: ``True``
|
- **Default**: ``True``
|
||||||
- **Description**: Whether to write the summary information file
|
- **Description**: Whether to write the summary information file.
|
||||||
|
See :ref:`disabling_file_generation`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
@@ -38,7 +42,8 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: string
|
- **Type**: string
|
||||||
- **Default**: ``llms.txt``
|
- **Default**: ``llms.txt``
|
||||||
- **Description**: Name of the summary information file
|
- **Description**: Name of the summary information file.
|
||||||
|
See :ref:`changing_filenames`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
@@ -47,6 +52,7 @@ Project Configuration Values
|
|||||||
- **Type**: list of strings
|
- **Type**: list of strings
|
||||||
- **Default**: ``[]`` (empty list)
|
- **Default**: ``[]`` (empty list)
|
||||||
- **Description**: List of custom directive names to process for path resolution.
|
- **Description**: List of custom directive names to process for path resolution.
|
||||||
|
See :ref:`path_resolution`.
|
||||||
|
|
||||||
.. versionadded:: 0.1.0
|
.. versionadded:: 0.1.0
|
||||||
|
|
||||||
@@ -55,14 +61,16 @@ Project Configuration Values
|
|||||||
- **Type**: string or ``None``
|
- **Type**: string or ``None``
|
||||||
- **Default**: ``None``
|
- **Default**: ``None``
|
||||||
- **Description**: Overrides the Sphinx project name as the heading in ``llms.txt``.
|
- **Description**: Overrides the Sphinx project name as the heading in ``llms.txt``.
|
||||||
|
See :ref:`custom_title`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
.. confval:: llms_txt_summary
|
.. confval:: llms_txt_summary
|
||||||
|
|
||||||
- **Type**: string or ``None``
|
- **Type**: string
|
||||||
- **Default**: ``None``
|
- **Default**: The first paragraph in the root document, else an empty string
|
||||||
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
|
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
|
||||||
|
See :ref:`custom_summary`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
@@ -70,14 +78,7 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: list of strings
|
- **Type**: list of strings
|
||||||
- **Default**: ``[]``
|
- **Default**: ``[]``
|
||||||
- **Description**: A list of pages to ignore (e.g., ``["page1", "page_with_*"]``).
|
- **Description**: A list of pages to ignore.
|
||||||
|
See :ref:`excluding_content`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.1
|
.. versionadded:: 0.2.1
|
||||||
|
|
||||||
.. confval:: llms_txt_rm_directives
|
|
||||||
|
|
||||||
- **Type**: boolean
|
|
||||||
- **Default**: ``False``
|
|
||||||
- **Description**: Whether to remove all directives from the output files.
|
|
||||||
|
|
||||||
.. versionadded:: 0.2.3
|
|
||||||
|
|||||||
@@ -1,6 +1,11 @@
|
|||||||
Getting Started
|
Getting Started
|
||||||
===============
|
===============
|
||||||
|
|
||||||
|
Demo
|
||||||
|
----
|
||||||
|
|
||||||
|
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
||||||
|
|
||||||
Installation
|
Installation
|
||||||
------------
|
------------
|
||||||
|
|
||||||
@@ -22,3 +27,24 @@ Add the extension to your Sphinx configuration (``conf.py``):
|
|||||||
]
|
]
|
||||||
|
|
||||||
Once added, the extension will automatically generate the LLMs.txt files during the build process.
|
Once added, the extension will automatically generate the LLMs.txt files during the build process.
|
||||||
|
|
||||||
|
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
|
||||||
|
|
||||||
|
How It Works
|
||||||
|
------------
|
||||||
|
|
||||||
|
During the Sphinx build process:
|
||||||
|
|
||||||
|
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
|
||||||
|
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
|
||||||
|
3. **Path Resolution**: Transforms relative paths in directives to full paths
|
||||||
|
4. **Output Generation**: Creates two optional files:
|
||||||
|
|
||||||
|
- ``llms.txt``: A concise summary of your documentation, in Markdown
|
||||||
|
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
|
||||||
|
|
||||||
|
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
|
||||||
|
|
||||||
|
|
||||||
|
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
||||||
|
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
||||||
|
|||||||
+11
-20
@@ -3,38 +3,29 @@ Sphinx llms.txt Generator
|
|||||||
|
|
||||||
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|
||||||
|
|
||||||
|PyPI version|
|
|PyPI version| |Downloads| |Parallel Safe| |GitHub Stars|
|
||||||
|
|
||||||
.. toctree::
|
.. toctree::
|
||||||
:maxdepth: 2
|
:maxdepth: 2
|
||||||
|
|
||||||
getting-started
|
getting-started
|
||||||
|
advanced-configuration
|
||||||
configuration-values
|
configuration-values
|
||||||
contributing
|
contributing
|
||||||
changelog
|
changelog
|
||||||
|
|
||||||
Features
|
|
||||||
--------
|
|
||||||
|
|
||||||
Sphinx LLMs.txt provides the following features:
|
|
||||||
|
|
||||||
- Creates ``llms.txt`` and ``llms-full.txt``
|
|
||||||
- Automatically add content from ``include`` directives
|
|
||||||
- Resolves relative paths in directives like ``image`` and ``figure`` to use full paths
|
|
||||||
- Ability to add list of custom directives with ``llms_txt_directives``
|
|
||||||
- Optionally, prepend a base URL using Sphinx's ``html_baseurl``
|
|
||||||
- Ability to exclude pages
|
|
||||||
|
|
||||||
Example
|
|
||||||
-------
|
|
||||||
|
|
||||||
You can see this Sphinx projects `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
|
||||||
|
|
||||||
|
|
||||||
.. _Sphinx: http://sphinx-doc.org/
|
.. _Sphinx: http://sphinx-doc.org/
|
||||||
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
|
||||||
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
|
||||||
|
|
||||||
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
||||||
:target: https://pypi.python.org/pypi/sphinx-llms-txt
|
:target: https://pypi.python.org/pypi/sphinx-llms-txt
|
||||||
:alt: Latest PyPi Version
|
:alt: Latest PyPi Version
|
||||||
|
.. |Downloads| image:: https://static.pepy.tech/badge/sphinx-llms-txt/month
|
||||||
|
:target: https://pepy.tech/project/sphinx-llms-txt
|
||||||
|
:alt: PyPi Downloads per month
|
||||||
|
.. |Parallel Safe| image:: https://img.shields.io/badge/parallel%20safe-true-brightgreen
|
||||||
|
:target: #
|
||||||
|
:alt: Parallel read/write safe
|
||||||
|
.. |GitHub Stars| image:: https://img.shields.io/github/stars/jdillard/sphinx-llms-txt?style=social
|
||||||
|
:target: https://github.com/jdillard/sphinx-llms-txt
|
||||||
|
:alt: GitHub Repository stars
|
||||||
|
|||||||
@@ -12,7 +12,7 @@ from .manager import LLMSFullManager
|
|||||||
from .processor import DocumentProcessor
|
from .processor import DocumentProcessor
|
||||||
from .writer import FileWriter
|
from .writer import FileWriter
|
||||||
|
|
||||||
__version__ = "0.2.2"
|
__version__ = "0.3.0"
|
||||||
|
|
||||||
# Export classes needed by tests
|
# Export classes needed by tests
|
||||||
__all__ = [
|
__all__ = [
|
||||||
@@ -25,9 +25,14 @@ __all__ = [
|
|||||||
# Global manager instance
|
# Global manager instance
|
||||||
_manager = LLMSFullManager()
|
_manager = LLMSFullManager()
|
||||||
|
|
||||||
|
# Store root document first paragraph
|
||||||
|
_root_first_paragraph = ""
|
||||||
|
|
||||||
|
|
||||||
def doctree_resolved(app: Sphinx, doctree, docname: str):
|
def doctree_resolved(app: Sphinx, doctree, docname: str):
|
||||||
"""Called when a docname has been resolved to a document."""
|
"""Called when a docname has been resolved to a document."""
|
||||||
|
global _root_first_paragraph
|
||||||
|
|
||||||
# Extract title from the document
|
# Extract title from the document
|
||||||
title = None
|
title = None
|
||||||
# findall() returns a generator, convert to list to check if it has elements
|
# findall() returns a generator, convert to list to check if it has elements
|
||||||
@@ -38,6 +43,14 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
|
|||||||
if title:
|
if title:
|
||||||
_manager.update_page_title(docname, title)
|
_manager.update_page_title(docname, title)
|
||||||
|
|
||||||
|
# Extract first paragraph from root document
|
||||||
|
if docname == app.config.master_doc:
|
||||||
|
for node in doctree.traverse(nodes.paragraph):
|
||||||
|
first_para = node.astext()
|
||||||
|
if first_para:
|
||||||
|
_root_first_paragraph = first_para
|
||||||
|
break
|
||||||
|
|
||||||
|
|
||||||
def build_finished(app: Sphinx, exception):
|
def build_finished(app: Sphinx, exception):
|
||||||
"""Called when the build is finished."""
|
"""Called when the build is finished."""
|
||||||
@@ -47,12 +60,17 @@ def build_finished(app: Sphinx, exception):
|
|||||||
_manager.set_master_doc(app.config.master_doc)
|
_manager.set_master_doc(app.config.master_doc)
|
||||||
_manager.set_app(app)
|
_manager.set_app(app)
|
||||||
|
|
||||||
|
# Get the summary - use configured value or extracted first paragraph
|
||||||
|
summary = app.config.llms_txt_summary
|
||||||
|
if summary is None:
|
||||||
|
summary = _root_first_paragraph
|
||||||
|
|
||||||
# Set up configuration
|
# Set up configuration
|
||||||
config = {
|
config = {
|
||||||
"llms_txt_file": app.config.llms_txt_file,
|
"llms_txt_file": app.config.llms_txt_file,
|
||||||
"llms_txt_filename": app.config.llms_txt_filename,
|
"llms_txt_filename": app.config.llms_txt_filename,
|
||||||
"llms_txt_title": app.config.llms_txt_title,
|
"llms_txt_title": app.config.llms_txt_title,
|
||||||
"llms_txt_summary": app.config.llms_txt_summary,
|
"llms_txt_summary": summary,
|
||||||
"llms_txt_full_file": app.config.llms_txt_full_file,
|
"llms_txt_full_file": app.config.llms_txt_full_file,
|
||||||
"llms_txt_full_filename": app.config.llms_txt_full_filename,
|
"llms_txt_full_filename": app.config.llms_txt_full_filename,
|
||||||
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
||||||
@@ -91,9 +109,10 @@ def setup(app: Sphinx) -> Dict[str, Any]:
|
|||||||
app.connect("doctree-resolved", doctree_resolved)
|
app.connect("doctree-resolved", doctree_resolved)
|
||||||
app.connect("build-finished", build_finished)
|
app.connect("build-finished", build_finished)
|
||||||
|
|
||||||
# Reset manager for each build
|
# Reset manager and root paragraph for each build
|
||||||
global _manager
|
global _manager, _root_first_paragraph
|
||||||
_manager = LLMSFullManager()
|
_manager = LLMSFullManager()
|
||||||
|
_root_first_paragraph = ""
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"version": __version__,
|
"version": __version__,
|
||||||
|
|||||||
+118
-25
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
import fnmatch
|
import fnmatch
|
||||||
from typing import Any, Dict, List
|
from typing import Any, Dict, List, Tuple
|
||||||
|
|
||||||
from sphinx.environment import BuildEnvironment
|
from sphinx.environment import BuildEnvironment
|
||||||
from sphinx.util import logging
|
from sphinx.util import logging
|
||||||
@@ -19,6 +19,7 @@ class DocumentCollector:
|
|||||||
self.master_doc: str = None
|
self.master_doc: str = None
|
||||||
self.env: BuildEnvironment = None
|
self.env: BuildEnvironment = None
|
||||||
self.config: Dict[str, Any] = {}
|
self.config: Dict[str, Any] = {}
|
||||||
|
self.app = None
|
||||||
|
|
||||||
def set_master_doc(self, master_doc: str):
|
def set_master_doc(self, master_doc: str):
|
||||||
"""Set the master document name."""
|
"""Set the master document name."""
|
||||||
@@ -37,8 +38,73 @@ class DocumentCollector:
|
|||||||
"""Set configuration options."""
|
"""Set configuration options."""
|
||||||
self.config = config
|
self.config = config
|
||||||
|
|
||||||
def get_page_order(self) -> List[str]:
|
def set_app(self, app):
|
||||||
"""Get the correct page order from the toctree structure."""
|
"""Set the Sphinx application reference."""
|
||||||
|
self.app = app
|
||||||
|
|
||||||
|
def _get_source_suffixes(self):
|
||||||
|
"""Get all valid source file suffixes from Sphinx configuration.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
|
||||||
|
"""
|
||||||
|
if not self.app:
|
||||||
|
return [".rst"] # Default fallback
|
||||||
|
|
||||||
|
source_suffix = self.app.config.source_suffix
|
||||||
|
|
||||||
|
if isinstance(source_suffix, dict):
|
||||||
|
return list(source_suffix.keys())
|
||||||
|
elif isinstance(source_suffix, list):
|
||||||
|
return source_suffix
|
||||||
|
else:
|
||||||
|
return [source_suffix] # String format
|
||||||
|
|
||||||
|
def _get_docname_suffix(self, docname: str, sources_dir) -> str:
|
||||||
|
"""
|
||||||
|
Determine the source suffix for a given docname by checking which
|
||||||
|
file exists.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
docname: The document name to check
|
||||||
|
sources_dir: Path to the _sources directory
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
The source suffix if found, or None if no matching file exists
|
||||||
|
"""
|
||||||
|
if not sources_dir or not sources_dir.exists():
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get the source link suffix from Sphinx config
|
||||||
|
source_link_suffix = ""
|
||||||
|
if self.app and hasattr(self.app.config, "html_sourcelink_suffix"):
|
||||||
|
source_link_suffix = self.app.config.html_sourcelink_suffix
|
||||||
|
# Handle empty string case specially
|
||||||
|
if source_link_suffix == "":
|
||||||
|
source_link_suffix = "" # Keep it empty
|
||||||
|
elif not source_link_suffix.startswith("."):
|
||||||
|
source_link_suffix = "." + source_link_suffix
|
||||||
|
|
||||||
|
# Get the source file suffixes from Sphinx config
|
||||||
|
source_suffixes = self._get_source_suffixes()
|
||||||
|
|
||||||
|
# Try to find the source file with any of the valid source suffixes
|
||||||
|
for src_suffix in source_suffixes:
|
||||||
|
candidate_file = sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
|
||||||
|
if candidate_file.exists():
|
||||||
|
return src_suffix
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
|
||||||
|
"""Get the correct page order from the toctree structure.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
sources_dir: Optional path to _sources directory for suffix detection
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List of tuples (docname, source_suffix) in toctree order
|
||||||
|
"""
|
||||||
if not self.env or not self.master_doc:
|
if not self.env or not self.master_doc:
|
||||||
return []
|
return []
|
||||||
|
|
||||||
@@ -52,9 +118,12 @@ class DocumentCollector:
|
|||||||
|
|
||||||
visited.add(docname)
|
visited.add(docname)
|
||||||
|
|
||||||
# Add the current document
|
# Add the current document with its suffix
|
||||||
if docname not in page_order:
|
if docname not in [doc for doc, _ in page_order]:
|
||||||
page_order.append(docname)
|
suffix = None
|
||||||
|
if sources_dir:
|
||||||
|
suffix = self._get_docname_suffix(docname, sources_dir)
|
||||||
|
page_order.append((docname, suffix))
|
||||||
|
|
||||||
# Check for toctree entries in this document
|
# Check for toctree entries in this document
|
||||||
try:
|
try:
|
||||||
@@ -65,20 +134,33 @@ class DocumentCollector:
|
|||||||
):
|
):
|
||||||
for child_docname in self.env.toctree_includes[docname]:
|
for child_docname in self.env.toctree_includes[docname]:
|
||||||
collect_from_toctree(child_docname)
|
collect_from_toctree(child_docname)
|
||||||
else:
|
# Try to use dependencies to find related documents
|
||||||
# Fallback: try to resolve and parse the toctree
|
elif (
|
||||||
toctree = self.env.get_and_resolve_toctree(docname, None)
|
hasattr(self.env, "dependencies")
|
||||||
if toctree:
|
and docname in self.env.dependencies
|
||||||
from docutils import nodes
|
):
|
||||||
|
# Extract the dependent documents from the dependencies dict
|
||||||
for node in list(toctree.findall(nodes.reference)):
|
for child_docname in self.env.dependencies[docname]:
|
||||||
if "refuri" in node.attributes:
|
# Only add documents actually in the document set
|
||||||
refuri = node.attributes["refuri"]
|
|
||||||
if refuri and refuri.endswith(".html"):
|
|
||||||
child_docname = refuri[:-5] # Remove .html
|
|
||||||
if (
|
if (
|
||||||
child_docname != docname
|
hasattr(self.env, "all_docs")
|
||||||
): # Avoid circular references
|
and child_docname in self.env.all_docs
|
||||||
|
):
|
||||||
|
collect_from_toctree(child_docname)
|
||||||
|
# Fallback to titles or other available references
|
||||||
|
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
|
||||||
|
# Get all document names
|
||||||
|
all_docnames = list(self.env.all_docs.keys())
|
||||||
|
|
||||||
|
# Look for documents that might be related (have similar paths)
|
||||||
|
current_prefix = "/".join(docname.split("/")[:-1])
|
||||||
|
if current_prefix:
|
||||||
|
for child_docname in all_docnames:
|
||||||
|
# Documents in the same directory might be related
|
||||||
|
if (
|
||||||
|
child_docname.startswith(current_prefix)
|
||||||
|
and child_docname != docname
|
||||||
|
):
|
||||||
collect_from_toctree(child_docname)
|
collect_from_toctree(child_docname)
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.debug(f"Could not get toctree for {docname}: {e}")
|
logger.debug(f"Could not get toctree for {docname}: {e}")
|
||||||
@@ -88,22 +170,33 @@ class DocumentCollector:
|
|||||||
|
|
||||||
# Add any remaining documents not in the toctree (sorted)
|
# Add any remaining documents not in the toctree (sorted)
|
||||||
if hasattr(self.env, "all_docs"):
|
if hasattr(self.env, "all_docs"):
|
||||||
|
processed_docnames = {doc for doc, _ in page_order}
|
||||||
remaining = sorted(
|
remaining = sorted(
|
||||||
[doc for doc in self.env.all_docs.keys() if doc not in page_order]
|
[
|
||||||
|
doc
|
||||||
|
for doc in self.env.all_docs.keys()
|
||||||
|
if doc not in processed_docnames
|
||||||
|
]
|
||||||
)
|
)
|
||||||
page_order.extend(remaining)
|
for docname in remaining:
|
||||||
|
suffix = None
|
||||||
|
if sources_dir:
|
||||||
|
suffix = self._get_docname_suffix(docname, sources_dir)
|
||||||
|
page_order.append((docname, suffix))
|
||||||
|
|
||||||
return page_order
|
return page_order
|
||||||
|
|
||||||
def filter_excluded_pages(self, page_order: List[str]) -> List[str]:
|
def filter_excluded_pages(
|
||||||
|
self, page_order: List[Tuple[str, str]]
|
||||||
|
) -> List[Tuple[str, str]]:
|
||||||
"""Filter out excluded pages from the page order."""
|
"""Filter out excluded pages from the page order."""
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
exclude_patterns = self.config.get("llms_txt_exclude")
|
||||||
if exclude_patterns:
|
if exclude_patterns:
|
||||||
return [
|
return [
|
||||||
page
|
(docname, suffix)
|
||||||
for page in page_order
|
for docname, suffix in page_order
|
||||||
if not any(
|
if not any(
|
||||||
self._match_exclude_pattern(page, pattern)
|
self._match_exclude_pattern(docname, pattern)
|
||||||
for pattern in exclude_patterns
|
for pattern in exclude_patterns
|
||||||
)
|
)
|
||||||
]
|
]
|
||||||
|
|||||||
+117
-92
@@ -56,6 +56,7 @@ class LLMSFullManager:
|
|||||||
def set_app(self, app: Sphinx):
|
def set_app(self, app: Sphinx):
|
||||||
"""Set the Sphinx application reference."""
|
"""Set the Sphinx application reference."""
|
||||||
self.app = app
|
self.app = app
|
||||||
|
self.collector.set_app(app)
|
||||||
if self.writer:
|
if self.writer:
|
||||||
self.writer.app = app
|
self.writer.app = app
|
||||||
|
|
||||||
@@ -69,23 +70,7 @@ class LLMSFullManager:
|
|||||||
self.processor = DocumentProcessor(self.config, srcdir)
|
self.processor = DocumentProcessor(self.config, srcdir)
|
||||||
self.writer = FileWriter(self.config, outdir, self.app)
|
self.writer = FileWriter(self.config, outdir, self.app)
|
||||||
|
|
||||||
# Get the correct page order
|
# Find sources directory first so we can pass it to get_page_order
|
||||||
page_order = self.collector.get_page_order()
|
|
||||||
|
|
||||||
if not page_order:
|
|
||||||
logger.warning(
|
|
||||||
"Could not determine page order, skipping llms-full creation"
|
|
||||||
)
|
|
||||||
return
|
|
||||||
|
|
||||||
# Apply exclusion filter if configured
|
|
||||||
page_order = self.collector.filter_excluded_pages(page_order)
|
|
||||||
|
|
||||||
# Determine output file name and location
|
|
||||||
output_filename = self.config.get("llms_txt_full_filename")
|
|
||||||
output_path = Path(outdir) / output_filename
|
|
||||||
|
|
||||||
# Find sources directory
|
|
||||||
sources_dir = None
|
sources_dir = None
|
||||||
possible_sources = [
|
possible_sources = [
|
||||||
Path(outdir) / "_sources",
|
Path(outdir) / "_sources",
|
||||||
@@ -104,14 +89,23 @@ class LLMSFullManager:
|
|||||||
)
|
)
|
||||||
return
|
return
|
||||||
|
|
||||||
# Collect all available source files
|
# Get the correct page order with source suffixes
|
||||||
txt_files = {}
|
page_order = self.collector.get_page_order(sources_dir)
|
||||||
for f in sources_dir.glob("**/*.txt"):
|
|
||||||
logger.debug(f"sphinx-llms-txt: Found source file: {f.stem} at {f}")
|
if not page_order:
|
||||||
txt_files[f.stem] = f
|
logger.warning(
|
||||||
|
"Could not determine page order, skipping llms-full creation"
|
||||||
|
)
|
||||||
|
return
|
||||||
|
|
||||||
|
# Apply exclusion filter if configured
|
||||||
|
page_order = self.collector.filter_excluded_pages(page_order)
|
||||||
|
|
||||||
|
# Determine output file name and location
|
||||||
|
output_filename = self.config.get("llms_txt_full_filename")
|
||||||
|
output_path = Path(outdir) / output_filename
|
||||||
|
|
||||||
# Log discovered files and page order
|
# Log discovered files and page order
|
||||||
logger.debug(f"sphinx-llms-txt: Found {len(txt_files)} source files")
|
|
||||||
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
|
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
|
||||||
|
|
||||||
# Log exclusion patterns
|
# Log exclusion patterns
|
||||||
@@ -119,33 +113,43 @@ class LLMSFullManager:
|
|||||||
if exclude_patterns:
|
if exclude_patterns:
|
||||||
logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}")
|
logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}")
|
||||||
|
|
||||||
# Create a mapping from docnames to actual file names
|
# Create a mapping from docnames to source files
|
||||||
docname_to_file = {}
|
docname_to_file = {}
|
||||||
|
|
||||||
# Try exact matches first
|
# Get the source link suffix from Sphinx config
|
||||||
for docname in page_order:
|
source_link_suffix = (
|
||||||
|
self.app.config.html_sourcelink_suffix if self.app else ".txt"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Handle empty string case specially
|
||||||
|
if source_link_suffix == "":
|
||||||
|
source_link_suffix = "" # Keep it empty
|
||||||
|
elif not source_link_suffix.startswith("."):
|
||||||
|
source_link_suffix = "." + source_link_suffix
|
||||||
|
|
||||||
|
# Process each (docname, suffix) in the page order
|
||||||
|
for docname, src_suffix in page_order:
|
||||||
# Skip excluded pages
|
# Skip excluded pages
|
||||||
if any(
|
if exclude_patterns and any(
|
||||||
self.collector._match_exclude_pattern(docname, pattern)
|
self.collector._match_exclude_pattern(docname, pattern)
|
||||||
for pattern in exclude_patterns
|
for pattern in exclude_patterns
|
||||||
):
|
):
|
||||||
continue
|
continue
|
||||||
|
|
||||||
if docname in txt_files:
|
# Build the source file path directly using the known suffix
|
||||||
docname_to_file[docname] = txt_files[docname]
|
if src_suffix:
|
||||||
|
source_file = sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
|
||||||
|
if source_file.exists():
|
||||||
|
docname_to_file[docname] = source_file
|
||||||
else:
|
else:
|
||||||
# Try with .rst extension
|
logger.warning(
|
||||||
if f"{docname}.rst" in txt_files:
|
f"sphinx-llms-txt: Source file not found for: {docname}."
|
||||||
docname_to_file[docname] = txt_files[f"{docname}.rst"]
|
f"Expected: {docname}{src_suffix}{source_link_suffix}"
|
||||||
# Try with .txt extension
|
)
|
||||||
elif f"{docname}.txt" in txt_files:
|
else:
|
||||||
docname_to_file[docname] = txt_files[f"{docname}.txt"]
|
logger.warning(
|
||||||
# Try with underscores instead of hyphens
|
f"sphinx-llms-txt: No source suffix determined for: {docname}"
|
||||||
elif docname.replace("-", "_") in txt_files:
|
)
|
||||||
docname_to_file[docname] = txt_files[docname.replace("-", "_")]
|
|
||||||
# Try with hyphens instead of underscores
|
|
||||||
elif docname.replace("_", "-") in txt_files:
|
|
||||||
docname_to_file[docname] = txt_files[docname.replace("_", "-")]
|
|
||||||
|
|
||||||
# Generate content
|
# Generate content
|
||||||
content_parts = []
|
content_parts = []
|
||||||
@@ -156,7 +160,7 @@ class LLMSFullManager:
|
|||||||
max_lines = self.config.get("llms_txt_full_max_size")
|
max_lines = self.config.get("llms_txt_full_max_size")
|
||||||
abort_due_to_max_lines = False
|
abort_due_to_max_lines = False
|
||||||
|
|
||||||
for docname in page_order:
|
for docname, _ in page_order:
|
||||||
if docname in docname_to_file:
|
if docname in docname_to_file:
|
||||||
file_path = docname_to_file[docname]
|
file_path = docname_to_file[docname]
|
||||||
content, line_count = self._read_source_file(file_path, docname)
|
content, line_count = self._read_source_file(file_path, docname)
|
||||||
@@ -190,65 +194,68 @@ class LLMSFullManager:
|
|||||||
added_files.add(file_path.stem)
|
added_files.add(file_path.stem)
|
||||||
total_line_count += line_count
|
total_line_count += line_count
|
||||||
else:
|
else:
|
||||||
logger.warning(f"sphinx-llm-txt: Source file not found for: {docname}")
|
logger.warning(
|
||||||
|
f"sphinx-llms-txt: Source file not found for: {docname}. Check that"
|
||||||
|
f" file exists at _sources/{docname}[suffix]{source_link_suffix}"
|
||||||
|
)
|
||||||
|
|
||||||
# Add any remaining files (in alphabetical order) if not aborted
|
# Add any remaining files (in alphabetical order) that aren't in the page order
|
||||||
if not abort_due_to_max_lines:
|
if not abort_due_to_max_lines:
|
||||||
# Apply the same exclusion filter to remaining files
|
# Get all source files in the _sources directory using configured suffixes
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
source_suffixes = self._get_source_suffixes()
|
||||||
|
all_source_files = []
|
||||||
|
for src_suffix in source_suffixes:
|
||||||
|
glob_pattern = f"**/*{src_suffix}{source_link_suffix}"
|
||||||
|
all_source_files.extend(sources_dir.glob(glob_pattern))
|
||||||
|
|
||||||
# Create a set of files to exclude based on their basename
|
processed_paths = set(file.resolve() for file in docname_to_file.values())
|
||||||
excluded_files = set()
|
|
||||||
for pattern in exclude_patterns:
|
|
||||||
if "*" not in pattern and "?" not in pattern:
|
|
||||||
# For exact patterns, add variants
|
|
||||||
excluded_files.add(pattern)
|
|
||||||
excluded_files.add(f"{pattern}.rst")
|
|
||||||
excluded_files.add(f"{pattern}.txt")
|
|
||||||
excluded_files.add(pattern.replace("-", "_"))
|
|
||||||
excluded_files.add(pattern.replace("_", "-"))
|
|
||||||
|
|
||||||
# Filter remaining files
|
# Find files that haven't been processed yet
|
||||||
remaining_files = sorted(
|
remaining_source_files = [
|
||||||
[
|
f for f in all_source_files if f.resolve() not in processed_paths
|
||||||
name
|
|
||||||
for name in txt_files
|
|
||||||
if name not in added_files
|
|
||||||
and name not in excluded_files
|
|
||||||
and not any(
|
|
||||||
self.collector._match_exclude_pattern(name, pattern)
|
|
||||||
for pattern in exclude_patterns
|
|
||||||
)
|
|
||||||
]
|
]
|
||||||
|
|
||||||
|
# Sort the remaining files for consistent ordering
|
||||||
|
remaining_source_files.sort()
|
||||||
|
|
||||||
|
if remaining_source_files:
|
||||||
|
logger.info(
|
||||||
|
f"Found {len(remaining_source_files)} additional files not in"
|
||||||
|
f" toctree"
|
||||||
)
|
)
|
||||||
if remaining_files:
|
|
||||||
logger.info(f"Adding remaining files: {remaining_files}")
|
for file_path in remaining_source_files:
|
||||||
for file_stem in remaining_files:
|
# Extract docname from path by removing the source and link suffixes
|
||||||
file_path = txt_files[file_stem]
|
rel_path = str(file_path.relative_to(sources_dir))
|
||||||
content, line_count = self._read_source_file(file_path, file_stem)
|
docname = None
|
||||||
|
|
||||||
|
# Try each source suffix to find which one this file uses
|
||||||
|
for src_suffix in source_suffixes:
|
||||||
|
combined_suffix = f"{src_suffix}{source_link_suffix}"
|
||||||
|
if rel_path.endswith(combined_suffix):
|
||||||
|
docname = rel_path[: -len(combined_suffix)] # Remove suffix
|
||||||
|
break
|
||||||
|
|
||||||
|
if docname is None:
|
||||||
|
continue
|
||||||
|
|
||||||
|
# Skip excluded docnames
|
||||||
|
if exclude_patterns and any(
|
||||||
|
self.collector._match_exclude_pattern(docname, pattern)
|
||||||
|
for pattern in exclude_patterns
|
||||||
|
):
|
||||||
|
logger.debug(f"sphinx-llms-txt: Skipping excluded file: {docname}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
# Read and process the file
|
||||||
|
content, line_count = self._read_source_file(file_path, docname)
|
||||||
|
|
||||||
# Check if adding this file would exceed the maximum line count
|
# Check if adding this file would exceed the maximum line count
|
||||||
if max_lines is not None and total_line_count + line_count > max_lines:
|
if max_lines is not None and total_line_count + line_count > max_lines:
|
||||||
break
|
break
|
||||||
|
|
||||||
# Double-check that this file should be included
|
if content:
|
||||||
should_include = True
|
logger.debug(f"sphinx-llms-txt: Adding remaining file: {docname}")
|
||||||
file_stem = file_path.stem
|
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
|
||||||
|
|
||||||
if exclude_patterns:
|
|
||||||
# Check stem against exclusion patterns
|
|
||||||
if any(
|
|
||||||
self.collector._match_exclude_pattern(file_stem, pattern)
|
|
||||||
for pattern in exclude_patterns
|
|
||||||
):
|
|
||||||
logger.debug(
|
|
||||||
"sphinx-llms-txt: Final exclusion check removed remaining"
|
|
||||||
f" file: {file_stem}"
|
|
||||||
)
|
|
||||||
should_include = False
|
|
||||||
|
|
||||||
if content and should_include:
|
|
||||||
content_parts.append(content)
|
content_parts.append(content)
|
||||||
total_line_count += line_count
|
total_line_count += line_count
|
||||||
|
|
||||||
@@ -258,7 +265,7 @@ class LLMSFullManager:
|
|||||||
max_lines is not None and total_line_count > max_lines
|
max_lines is not None and total_line_count > max_lines
|
||||||
):
|
):
|
||||||
logger.warning(
|
logger.warning(
|
||||||
f"sphinx-llm-txt: Max line limit ({max_lines}) exceeded:"
|
f"sphinx-llms-txt: Max line limit ({max_lines}) exceeded:"
|
||||||
f" {total_line_count} > {max_lines}. "
|
f" {total_line_count} > {max_lines}. "
|
||||||
f"Not creating llms-full.txt file."
|
f"Not creating llms-full.txt file."
|
||||||
)
|
)
|
||||||
@@ -325,5 +332,23 @@ class LLMSFullManager:
|
|||||||
return content_str, line_count + 1
|
return content_str, line_count + 1
|
||||||
|
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.error(f"sphinx-llm-txt: Error reading source file {file_path}: {e}")
|
logger.error(f"sphinx-llms-txt: Error reading source file {file_path}: {e}")
|
||||||
return "", 0
|
return "", 0
|
||||||
|
|
||||||
|
def _get_source_suffixes(self):
|
||||||
|
"""Get all valid source file suffixes from Sphinx configuration.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
|
||||||
|
"""
|
||||||
|
if not self.app:
|
||||||
|
return [".rst"] # Default fallback
|
||||||
|
|
||||||
|
source_suffix = self.app.config.source_suffix
|
||||||
|
|
||||||
|
if isinstance(source_suffix, dict):
|
||||||
|
return list(source_suffix.keys())
|
||||||
|
elif isinstance(source_suffix, list):
|
||||||
|
return source_suffix
|
||||||
|
else:
|
||||||
|
return [source_suffix] # String format
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Dict, List
|
from typing import Any, Dict, List, Tuple, Union
|
||||||
|
|
||||||
from sphinx.application import Sphinx
|
from sphinx.application import Sphinx
|
||||||
from sphinx.util import logging
|
from sphinx.util import logging
|
||||||
@@ -42,19 +42,19 @@ class FileWriter:
|
|||||||
)
|
)
|
||||||
return True
|
return True
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.error(f"sphinx-llm-txt: Error writing combined sources file: {e}")
|
logger.error(f"sphinx-llms-txt: Error writing combined sources file: {e}")
|
||||||
return False
|
return False
|
||||||
|
|
||||||
def write_verbose_info_to_file(
|
def write_verbose_info_to_file(
|
||||||
self,
|
self,
|
||||||
page_order: List[str],
|
page_order: Union[List[str], List[Tuple[str, str]]],
|
||||||
page_titles: Dict[str, str],
|
page_titles: Dict[str, str],
|
||||||
total_line_count: int = 0,
|
total_line_count: int = 0,
|
||||||
) -> bool:
|
) -> bool:
|
||||||
"""Write summary information to the llms.txt file.
|
"""Write summary information to the llms.txt file.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
page_order: Ordered list of document names
|
page_order: Ordered list of document names or (docname, suffix) tuples
|
||||||
page_titles: Dictionary mapping docnames to titles
|
page_titles: Dictionary mapping docnames to titles
|
||||||
total_line_count: Total number of lines in the combined content
|
total_line_count: Total number of lines in the combined content
|
||||||
|
|
||||||
@@ -88,6 +88,8 @@ class FileWriter:
|
|||||||
if description:
|
if description:
|
||||||
# Trim leading and trailing whitespace
|
# Trim leading and trailing whitespace
|
||||||
description = description.strip()
|
description = description.strip()
|
||||||
|
if description:
|
||||||
|
# Only add blockquote if description is not empty
|
||||||
# Replace newlines with newline + blockquote marker to maintain
|
# Replace newlines with newline + blockquote marker to maintain
|
||||||
# blockquote formatting
|
# blockquote formatting
|
||||||
description = description.replace("\n", "\n> ")
|
description = description.replace("\n", "\n> ")
|
||||||
@@ -100,7 +102,12 @@ class FileWriter:
|
|||||||
if not base_url.endswith("/"):
|
if not base_url.endswith("/"):
|
||||||
base_url += "/"
|
base_url += "/"
|
||||||
|
|
||||||
for docname in page_order:
|
for item in page_order:
|
||||||
|
# Handle both old format (str) and new format (tuple)
|
||||||
|
if isinstance(item, tuple):
|
||||||
|
docname, _ = item
|
||||||
|
else:
|
||||||
|
docname = item
|
||||||
title = page_titles.get(docname, docname)
|
title = page_titles.get(docname, docname)
|
||||||
f.write(f"- [{title}]({base_url}{docname}.html)\n")
|
f.write(f"- [{title}]({base_url}{docname}.html)\n")
|
||||||
|
|
||||||
|
|||||||
@@ -334,3 +334,477 @@ def test_write_verbose_info_with_baseurl(tmp_path):
|
|||||||
|
|
||||||
assert "- [Home Page](https://example.org/index.html)" in content
|
assert "- [Home Page](https://example.org/index.html)" in content
|
||||||
assert "- [About Us](https://example.org/about.html)" in content
|
assert "- [About Us](https://example.org/about.html)" in content
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_source_suffixes_with_dict():
|
||||||
|
"""Test _get_source_suffixes method with dict source_suffix."""
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with dict source_suffix
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
source_suffix = {".rst": None, ".md": None, ".txt": None}
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
|
||||||
|
suffixes = manager._get_source_suffixes()
|
||||||
|
assert set(suffixes) == {".rst", ".md", ".txt"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_source_suffixes_with_list():
|
||||||
|
"""Test _get_source_suffixes method with list source_suffix."""
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with list source_suffix
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
source_suffix = [".rst", ".md"]
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
|
||||||
|
suffixes = manager._get_source_suffixes()
|
||||||
|
assert suffixes == [".rst", ".md"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_source_suffixes_with_string():
|
||||||
|
"""Test _get_source_suffixes method with string source_suffix."""
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with string source_suffix
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
source_suffix = ".rst"
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
|
||||||
|
suffixes = manager._get_source_suffixes()
|
||||||
|
assert suffixes == [".rst"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_source_suffixes_no_app():
|
||||||
|
"""Test _get_source_suffixes method with no app set."""
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
|
||||||
|
suffixes = manager._get_source_suffixes()
|
||||||
|
assert suffixes == [".rst"] # Default fallback
|
||||||
|
|
||||||
|
|
||||||
|
def test_html_sourcelink_suffix_default():
|
||||||
|
"""Test html_sourcelink_suffix defaults to .txt when no app is set."""
|
||||||
|
import tempfile
|
||||||
|
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_config(
|
||||||
|
{
|
||||||
|
"llms_txt_full_filename": "test.txt",
|
||||||
|
"llms_txt_exclude": [],
|
||||||
|
"llms_txt_directives": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
# Create a temporary directory structure
|
||||||
|
with tempfile.TemporaryDirectory() as tmpdir:
|
||||||
|
outdir = f"{tmpdir}/build"
|
||||||
|
srcdir = f"{tmpdir}/source"
|
||||||
|
sources_dir = f"{outdir}/_sources"
|
||||||
|
|
||||||
|
# Create directories
|
||||||
|
import os
|
||||||
|
|
||||||
|
os.makedirs(sources_dir, exist_ok=True)
|
||||||
|
os.makedirs(srcdir, exist_ok=True)
|
||||||
|
|
||||||
|
# Create a test source file with default .txt suffix
|
||||||
|
test_file = f"{sources_dir}/index.rst.txt"
|
||||||
|
with open(test_file, "w") as f:
|
||||||
|
f.write("Test content")
|
||||||
|
|
||||||
|
# Mock env with minimal required attributes
|
||||||
|
class MockEnv:
|
||||||
|
all_docs = {"index": None}
|
||||||
|
titles = {
|
||||||
|
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
|
||||||
|
}
|
||||||
|
toctree_includes = {}
|
||||||
|
|
||||||
|
manager.set_env(MockEnv())
|
||||||
|
manager.set_master_doc("index")
|
||||||
|
|
||||||
|
# Test that it uses .txt as the default suffix
|
||||||
|
manager.combine_sources(outdir, srcdir)
|
||||||
|
|
||||||
|
# Verify the file was found and processed (check if output file exists)
|
||||||
|
output_file = f"{outdir}/test.txt"
|
||||||
|
assert os.path.exists(output_file)
|
||||||
|
|
||||||
|
|
||||||
|
def test_html_sourcelink_suffix_custom():
|
||||||
|
"""Test html_sourcelink_suffix uses custom value from Sphinx config."""
|
||||||
|
import tempfile
|
||||||
|
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with custom html_sourcelink_suffix
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
html_sourcelink_suffix = "source"
|
||||||
|
source_suffix = ".rst"
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
manager.set_config(
|
||||||
|
{
|
||||||
|
"llms_txt_full_filename": "test.txt",
|
||||||
|
"llms_txt_exclude": [],
|
||||||
|
"llms_txt_directives": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
# Create a temporary directory structure
|
||||||
|
with tempfile.TemporaryDirectory() as tmpdir:
|
||||||
|
outdir = f"{tmpdir}/build"
|
||||||
|
srcdir = f"{tmpdir}/source"
|
||||||
|
sources_dir = f"{outdir}/_sources"
|
||||||
|
|
||||||
|
# Create directories
|
||||||
|
import os
|
||||||
|
|
||||||
|
os.makedirs(sources_dir, exist_ok=True)
|
||||||
|
os.makedirs(srcdir, exist_ok=True)
|
||||||
|
|
||||||
|
# Create a test source file with custom .source suffix
|
||||||
|
test_file = f"{sources_dir}/index.rst.source"
|
||||||
|
with open(test_file, "w") as f:
|
||||||
|
f.write("Test content")
|
||||||
|
|
||||||
|
# Mock env with minimal required attributes
|
||||||
|
class MockEnv:
|
||||||
|
all_docs = {"index": None}
|
||||||
|
titles = {
|
||||||
|
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
|
||||||
|
}
|
||||||
|
toctree_includes = {}
|
||||||
|
|
||||||
|
manager.set_env(MockEnv())
|
||||||
|
manager.set_master_doc("index")
|
||||||
|
|
||||||
|
# Test that it uses .source as the custom suffix
|
||||||
|
manager.combine_sources(outdir, srcdir)
|
||||||
|
|
||||||
|
# Verify the file was found and processed
|
||||||
|
output_file = f"{outdir}/test.txt"
|
||||||
|
assert os.path.exists(output_file)
|
||||||
|
|
||||||
|
|
||||||
|
def test_html_sourcelink_suffix_with_dot():
|
||||||
|
"""Test html_sourcelink_suffix adds dot if missing."""
|
||||||
|
import tempfile
|
||||||
|
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with html_sourcelink_suffix without leading dot
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
html_sourcelink_suffix = "src" # No leading dot
|
||||||
|
source_suffix = ".rst"
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
manager.set_config(
|
||||||
|
{
|
||||||
|
"llms_txt_full_filename": "test.txt",
|
||||||
|
"llms_txt_exclude": [],
|
||||||
|
"llms_txt_directives": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
# Create a temporary directory structure
|
||||||
|
with tempfile.TemporaryDirectory() as tmpdir:
|
||||||
|
outdir = f"{tmpdir}/build"
|
||||||
|
srcdir = f"{tmpdir}/source"
|
||||||
|
sources_dir = f"{outdir}/_sources"
|
||||||
|
|
||||||
|
# Create directories
|
||||||
|
import os
|
||||||
|
|
||||||
|
os.makedirs(sources_dir, exist_ok=True)
|
||||||
|
os.makedirs(srcdir, exist_ok=True)
|
||||||
|
|
||||||
|
# Create a test source file with .src suffix (dot should be added automatically)
|
||||||
|
test_file = f"{sources_dir}/index.rst.src"
|
||||||
|
with open(test_file, "w") as f:
|
||||||
|
f.write("Test content")
|
||||||
|
|
||||||
|
# Mock env with minimal required attributes
|
||||||
|
class MockEnv:
|
||||||
|
all_docs = {"index": None}
|
||||||
|
titles = {
|
||||||
|
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
|
||||||
|
}
|
||||||
|
toctree_includes = {}
|
||||||
|
|
||||||
|
manager.set_env(MockEnv())
|
||||||
|
manager.set_master_doc("index")
|
||||||
|
|
||||||
|
# Test that it adds the dot and finds the file
|
||||||
|
manager.combine_sources(outdir, srcdir)
|
||||||
|
|
||||||
|
# Verify the file was found and processed
|
||||||
|
output_file = f"{outdir}/test.txt"
|
||||||
|
assert os.path.exists(output_file)
|
||||||
|
|
||||||
|
|
||||||
|
def test_mixed_source_file_formats():
|
||||||
|
"""Test handling of mixed source file formats (.rst, .md, .txt)."""
|
||||||
|
import tempfile
|
||||||
|
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with multiple source suffixes
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
html_sourcelink_suffix = ".txt"
|
||||||
|
source_suffix = {".rst": None, ".md": None, ".txt": None}
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
manager.set_config(
|
||||||
|
{
|
||||||
|
"llms_txt_full_filename": "test.txt",
|
||||||
|
"llms_txt_exclude": [],
|
||||||
|
"llms_txt_directives": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
# Create a temporary directory structure
|
||||||
|
with tempfile.TemporaryDirectory() as tmpdir:
|
||||||
|
outdir = f"{tmpdir}/build"
|
||||||
|
srcdir = f"{tmpdir}/source"
|
||||||
|
sources_dir = f"{outdir}/_sources"
|
||||||
|
|
||||||
|
# Create directories
|
||||||
|
import os
|
||||||
|
|
||||||
|
os.makedirs(sources_dir, exist_ok=True)
|
||||||
|
os.makedirs(srcdir, exist_ok=True)
|
||||||
|
|
||||||
|
# Create test source files with different formats
|
||||||
|
files_to_create = [
|
||||||
|
f"{sources_dir}/page1.rst.txt",
|
||||||
|
f"{sources_dir}/page2.md.txt",
|
||||||
|
f"{sources_dir}/page3.txt.txt",
|
||||||
|
]
|
||||||
|
|
||||||
|
for test_file in files_to_create:
|
||||||
|
with open(test_file, "w") as f:
|
||||||
|
f.write(f"Content for {os.path.basename(test_file)}")
|
||||||
|
|
||||||
|
# Mock env with all documents
|
||||||
|
class MockEnv:
|
||||||
|
all_docs = {"page1": None, "page2": None, "page3": None}
|
||||||
|
titles = {
|
||||||
|
"page1": type("TitleNode", (), {"astext": lambda: "Page 1"})(),
|
||||||
|
"page2": type("TitleNode", (), {"astext": lambda: "Page 2"})(),
|
||||||
|
"page3": type("TitleNode", (), {"astext": lambda: "Page 3"})(),
|
||||||
|
}
|
||||||
|
toctree_includes = {}
|
||||||
|
|
||||||
|
manager.set_env(MockEnv())
|
||||||
|
manager.set_master_doc("page1")
|
||||||
|
|
||||||
|
# Test that all file formats are found and processed
|
||||||
|
manager.combine_sources(outdir, srcdir)
|
||||||
|
|
||||||
|
# Verify the output file was created and contains content from all formats
|
||||||
|
output_file = f"{outdir}/test.txt"
|
||||||
|
assert os.path.exists(output_file)
|
||||||
|
|
||||||
|
with open(output_file, "r") as f:
|
||||||
|
content = f.read()
|
||||||
|
|
||||||
|
# Should contain content from all three files
|
||||||
|
assert "Content for page1.rst.txt" in content
|
||||||
|
assert "Content for page2.md.txt" in content
|
||||||
|
assert "Content for page3.txt.txt" in content
|
||||||
|
|
||||||
|
|
||||||
|
def test_source_suffix_detection_priority():
|
||||||
|
"""Test source suffix detection tries formats in correct order for docnames."""
|
||||||
|
import tempfile
|
||||||
|
|
||||||
|
from sphinx_llms_txt.manager import LLMSFullManager
|
||||||
|
|
||||||
|
# Mock Sphinx app with ordered source suffixes
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
html_sourcelink_suffix = ".txt"
|
||||||
|
source_suffix = [".rst", ".md"] # rst has priority over md
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
manager.set_app(MockApp())
|
||||||
|
manager.set_config(
|
||||||
|
{
|
||||||
|
"llms_txt_full_filename": "test.txt",
|
||||||
|
"llms_txt_exclude": [],
|
||||||
|
"llms_txt_directives": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
# Create a temporary directory structure
|
||||||
|
with tempfile.TemporaryDirectory() as tmpdir:
|
||||||
|
outdir = f"{tmpdir}/build"
|
||||||
|
srcdir = f"{tmpdir}/source"
|
||||||
|
sources_dir = f"{outdir}/_sources"
|
||||||
|
|
||||||
|
# Create directories
|
||||||
|
import os
|
||||||
|
|
||||||
|
os.makedirs(sources_dir, exist_ok=True)
|
||||||
|
os.makedirs(srcdir, exist_ok=True)
|
||||||
|
|
||||||
|
# Create both .rst and .md versions of the same document
|
||||||
|
# Only create files for the specific docname "index"
|
||||||
|
rst_file = f"{sources_dir}/index.rst.txt"
|
||||||
|
md_file = f"{sources_dir}/index.md.txt"
|
||||||
|
|
||||||
|
with open(rst_file, "w") as f:
|
||||||
|
f.write("RST content for index")
|
||||||
|
|
||||||
|
with open(md_file, "w") as f:
|
||||||
|
f.write("Markdown content for index")
|
||||||
|
|
||||||
|
# Mock env with only the index document
|
||||||
|
class MockEnv:
|
||||||
|
all_docs = {"index": None}
|
||||||
|
titles = {
|
||||||
|
"index": type("TitleNode", (), {"astext": lambda: "Index Page"})()
|
||||||
|
}
|
||||||
|
toctree_includes = {"index": []}
|
||||||
|
|
||||||
|
manager.set_env(MockEnv())
|
||||||
|
manager.set_master_doc("index")
|
||||||
|
|
||||||
|
# Test the priority behavior
|
||||||
|
manager.combine_sources(outdir, srcdir)
|
||||||
|
|
||||||
|
# Check that output file was created
|
||||||
|
output_file = f"{outdir}/test.txt"
|
||||||
|
assert os.path.exists(output_file)
|
||||||
|
|
||||||
|
with open(output_file, "r") as f:
|
||||||
|
content = f.read()
|
||||||
|
|
||||||
|
# The system should prefer RST over MD for the "index" docname
|
||||||
|
# But since both files exist and the second phase adds remaining files,
|
||||||
|
# both will be included. The test verifies that RST appears first
|
||||||
|
# (indicating it was found first in the priority order)
|
||||||
|
assert "RST content for index" in content
|
||||||
|
|
||||||
|
# Find positions to verify order
|
||||||
|
rst_pos = content.find("RST content for index")
|
||||||
|
md_pos = content.find("Markdown content for index")
|
||||||
|
|
||||||
|
# RST should come before MD (due to priority in toctree processing)
|
||||||
|
assert rst_pos < md_pos, "RST content should appear before MD content"
|
||||||
|
|
||||||
|
|
||||||
|
def test_summary_default_uses_first_paragraph():
|
||||||
|
"""
|
||||||
|
Test that summary defaults to first paragraph of root document when not configured.
|
||||||
|
"""
|
||||||
|
from docutils import nodes
|
||||||
|
from docutils.frontend import OptionParser
|
||||||
|
from docutils.parsers.rst import Parser
|
||||||
|
from docutils.utils import new_document
|
||||||
|
|
||||||
|
from sphinx_llms_txt import build_finished, doctree_resolved
|
||||||
|
|
||||||
|
# Create a proper document with settings
|
||||||
|
settings = OptionParser(components=(Parser,)).get_default_values()
|
||||||
|
doctree = new_document("<rst-doc>", settings)
|
||||||
|
|
||||||
|
title = nodes.title(text="Test Title")
|
||||||
|
paragraph = nodes.paragraph(
|
||||||
|
text="This is the first paragraph that should be used as summary."
|
||||||
|
)
|
||||||
|
doctree.append(title)
|
||||||
|
doctree.append(paragraph)
|
||||||
|
|
||||||
|
# Mock Sphinx app
|
||||||
|
class MockApp:
|
||||||
|
class Config:
|
||||||
|
master_doc = "index"
|
||||||
|
llms_txt_summary = None # Not configured
|
||||||
|
llms_txt_file = True
|
||||||
|
llms_txt_filename = "llms.txt"
|
||||||
|
llms_txt_title = None
|
||||||
|
llms_txt_full_file = True
|
||||||
|
llms_txt_full_filename = "llms-full.txt"
|
||||||
|
llms_txt_full_max_size = None
|
||||||
|
llms_txt_directives = []
|
||||||
|
llms_txt_exclude = []
|
||||||
|
html_baseurl = ""
|
||||||
|
|
||||||
|
config = Config()
|
||||||
|
outdir = "/tmp/build"
|
||||||
|
srcdir = "/tmp/source"
|
||||||
|
|
||||||
|
class Env:
|
||||||
|
titles = {
|
||||||
|
"index": type("TitleNode", (), {"astext": lambda self: "Test Title"})()
|
||||||
|
}
|
||||||
|
|
||||||
|
env = Env()
|
||||||
|
|
||||||
|
app = MockApp()
|
||||||
|
|
||||||
|
# Reset the global state
|
||||||
|
import sphinx_llms_txt
|
||||||
|
|
||||||
|
sphinx_llms_txt._root_first_paragraph = ""
|
||||||
|
|
||||||
|
# Call doctree_resolved to extract the first paragraph
|
||||||
|
doctree_resolved(app, doctree, "index")
|
||||||
|
|
||||||
|
# Verify the first paragraph was extracted
|
||||||
|
assert (
|
||||||
|
sphinx_llms_txt._root_first_paragraph
|
||||||
|
== "This is the first paragraph that should be used as summary."
|
||||||
|
)
|
||||||
|
|
||||||
|
# Mock the manager methods to avoid actual file operations
|
||||||
|
original_combine_sources = sphinx_llms_txt._manager.combine_sources
|
||||||
|
sphinx_llms_txt._manager.combine_sources = lambda outdir, srcdir: None
|
||||||
|
|
||||||
|
# Call build_finished and verify the summary is set correctly
|
||||||
|
build_finished(app, None)
|
||||||
|
|
||||||
|
# Check that the summary was properly configured
|
||||||
|
assert (
|
||||||
|
sphinx_llms_txt._manager.config["llms_txt_summary"]
|
||||||
|
== "This is the first paragraph that should be used as summary."
|
||||||
|
)
|
||||||
|
|
||||||
|
# Restore original method
|
||||||
|
sphinx_llms_txt._manager.combine_sources = original_combine_sources
|
||||||
|
|||||||
Reference in New Issue
Block a user