Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6962ea1f80 | ||
|
|
7e390546ba | ||
|
|
ad987f77f0 | ||
|
|
bb75a8ba77 | ||
|
|
8da5aa9052 | ||
|
|
75380589e1 | ||
|
|
19c224c199 | ||
|
|
ebd0e13594 | ||
|
|
b4cab5ab52 | ||
|
|
c24f92031c |
@@ -18,6 +18,13 @@ repos:
|
|||||||
hooks:
|
hooks:
|
||||||
- id: flake8
|
- id: flake8
|
||||||
|
|
||||||
|
- repo: https://github.com/pre-commit/mirrors-mypy
|
||||||
|
rev: v1.11.2
|
||||||
|
hooks:
|
||||||
|
- id: mypy
|
||||||
|
files: ^sphinx_llms_txt/
|
||||||
|
additional_dependencies: [types-docutils]
|
||||||
|
|
||||||
- repo: https://github.com/sphinx-contrib/sphinx-lint
|
- repo: https://github.com/sphinx-contrib/sphinx-lint
|
||||||
rev: v1.0.0
|
rev: v1.0.0
|
||||||
hooks:
|
hooks:
|
||||||
|
|||||||
@@ -1,6 +1,20 @@
|
|||||||
Changelog
|
Changelog
|
||||||
=========
|
=========
|
||||||
|
|
||||||
|
0.5.0
|
||||||
|
-----
|
||||||
|
|
||||||
|
- Add :ref:`block_level_ignore` and :ref:`page_level_ignore`
|
||||||
|
`#33 <https://github.com/jdillard/sphinx-llms-txt/pull/33>`_
|
||||||
|
- Add :confval:`llms_txt_full_size_policy` configuration option to control behavior when :confval:`llms_txt_full_max_size` is exceeded.
|
||||||
|
`#35 <https://github.com/jdillard/sphinx-llms-txt/pull/35>`_
|
||||||
|
|
||||||
|
0.4.1
|
||||||
|
-----
|
||||||
|
|
||||||
|
- Fix include paths and spacing
|
||||||
|
`#31 <https://github.com/jdillard/sphinx-llms-txt/pull/31>`_
|
||||||
|
|
||||||
0.4.0
|
0.4.0
|
||||||
-----
|
-----
|
||||||
|
|
||||||
|
|||||||
@@ -11,6 +11,10 @@ A Sphinx extension that generates a summary `llms.txt` file and a single combine
|
|||||||
|
|
||||||
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
|
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
|
||||||
|
|
||||||
|
## Contributing
|
||||||
|
|
||||||
|
Pull Requests welcome! See [Contributing](https://sphinx-llms-txt.readthedocs.io/en/latest/contributing.html) for instructions on how best to contribute.
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
MIT License - see LICENSE file for details.
|
MIT License - see LICENSE file for details.
|
||||||
|
|||||||
@@ -76,15 +76,26 @@ Handling Large Documentation
|
|||||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
|
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
|
||||||
You can set a maximum line count:
|
You can set a maximum line count and control what happens when that limit is exceeded:
|
||||||
|
|
||||||
.. code-block:: python
|
.. code-block:: python
|
||||||
|
|
||||||
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
|
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
|
||||||
|
llms_txt_full_size_policy = "warn_skip" # Default behavior
|
||||||
|
|
||||||
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
|
The ``llms_txt_full_size_policy`` setting controls both the log level and action taken when the size limit is exceeded.
|
||||||
|
It uses the format ``"<loglevel>_<action>"``:
|
||||||
|
|
||||||
.. tip:: Use :ref:`excluding_content` to remove less relevant pages.
|
**Log levels:**
|
||||||
|
- ``warn``: Log as a warning (default)
|
||||||
|
- ``info``: Log as informational message
|
||||||
|
|
||||||
|
**Actions:**
|
||||||
|
- ``skip``: Don't create the file (default)
|
||||||
|
- ``keep``: Create the file anyway, ignoring the size limit
|
||||||
|
- ``note``: Create a placeholder file explaining why the full file wasn't generated
|
||||||
|
|
||||||
|
.. tip:: Use :ref:`excluding_content` to remove less relevant pages and reduce the file size.
|
||||||
|
|
||||||
.. _custom_directive_handling:
|
.. _custom_directive_handling:
|
||||||
|
|
||||||
@@ -113,6 +124,13 @@ This ensures that paths in your custom directives are properly resolved in the g
|
|||||||
Excluding Content
|
Excluding Content
|
||||||
^^^^^^^^^^^^^^^^^
|
^^^^^^^^^^^^^^^^^
|
||||||
|
|
||||||
|
There are several ways to exclude content from the generated ``llms-full.txt`` file:
|
||||||
|
|
||||||
|
.. _global_exclusion:
|
||||||
|
|
||||||
|
Global Page Exclusion
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
You can exclude specific pages from being included in the generated files:
|
You can exclude specific pages from being included in the generated files:
|
||||||
|
|
||||||
.. code-block:: python
|
.. code-block:: python
|
||||||
@@ -124,6 +142,67 @@ You can exclude specific pages from being included in the generated files:
|
|||||||
]
|
]
|
||||||
|
|
||||||
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
|
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
|
||||||
|
It can also be used to reduce the size of llms-full.txt.
|
||||||
|
|
||||||
|
.. _page_level_ignore:
|
||||||
|
|
||||||
|
Page-Level Ignore Metadata
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
You can exclude individual pages by adding metadata at the top of any reStructuredText file:
|
||||||
|
|
||||||
|
.. code-block:: restructuredtext
|
||||||
|
|
||||||
|
:llms-txt-ignore: true
|
||||||
|
|
||||||
|
Page Title
|
||||||
|
==========
|
||||||
|
|
||||||
|
This entire page will be excluded from llms-full.txt
|
||||||
|
|
||||||
|
When this metadata is present, the entire page is skipped during processing.
|
||||||
|
|
||||||
|
.. _block_level_ignore:
|
||||||
|
|
||||||
|
Block-Level Ignore Directives
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
|
||||||
|
You can exclude specific sections within a page using ignore directives:
|
||||||
|
|
||||||
|
.. code-block:: restructuredtext
|
||||||
|
|
||||||
|
Page Title
|
||||||
|
==========
|
||||||
|
|
||||||
|
This content will be included in llms-full.txt.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
This content will be excluded from llms-full.txt.
|
||||||
|
|
||||||
|
Section To Ignore
|
||||||
|
-----------------
|
||||||
|
|
||||||
|
This entire section and any nested content will be ignored.
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
# This code block will also be ignored
|
||||||
|
def ignored_function():
|
||||||
|
pass
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
This content will be included again.
|
||||||
|
|
||||||
|
Block-level ignores can be useful for:
|
||||||
|
|
||||||
|
- Removing internal notes or TODOs
|
||||||
|
- Hiding implementation details while keeping user-facing documentation
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
- Multiple ignore blocks can be used within the same file
|
||||||
|
- Ignore directives work with any indentation level
|
||||||
|
|
||||||
.. _including_code_files:
|
.. _including_code_files:
|
||||||
|
|
||||||
@@ -198,6 +277,7 @@ Here's a complete example showing multiple :doc:`configuration-values`:
|
|||||||
llms_txt_filename = "ai-summary.txt"
|
llms_txt_filename = "ai-summary.txt"
|
||||||
llms_txt_full_filename = "ai-full-docs.txt"
|
llms_txt_full_filename = "ai-full-docs.txt"
|
||||||
llms_txt_full_max_size = 50000
|
llms_txt_full_max_size = 50000
|
||||||
|
llms_txt_full_size_policy = "warn_note"
|
||||||
|
|
||||||
# Content customization
|
# Content customization
|
||||||
llms_txt_title = "Project Documentation for AI Assistants"
|
llms_txt_title = "Project Documentation for AI Assistants"
|
||||||
|
|||||||
@@ -24,11 +24,22 @@ Project Configuration Values
|
|||||||
- **Type**: integer or ``None``
|
- **Type**: integer or ``None``
|
||||||
- **Default**: ``None`` (no limit)
|
- **Default**: ``None`` (no limit)
|
||||||
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
|
||||||
If exceeded, the file is skipped and a warning is shown, but the build still completes.
|
Behavior when exceeded is controlled by :confval:`llms_txt_full_size_policy`.
|
||||||
See :ref:`handling_large_documentation`.
|
See :ref:`handling_large_documentation`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.0
|
.. versionadded:: 0.2.0
|
||||||
|
|
||||||
|
.. confval:: llms_txt_full_size_policy
|
||||||
|
|
||||||
|
- **Type**: string
|
||||||
|
- **Default**: ``'warn_skip'``
|
||||||
|
- **Description**: Controls what happens when :confval:`llms_txt_full_max_size` is exceeded.
|
||||||
|
Format is ``<loglevel>_<action>``. Log levels: ``warn``, ``info``.
|
||||||
|
Actions: ``skip``, ``keep``, ``note``.
|
||||||
|
See :ref:`handling_large_documentation`.
|
||||||
|
|
||||||
|
.. versionadded:: 0.5.0
|
||||||
|
|
||||||
.. confval:: llms_txt_file
|
.. confval:: llms_txt_file
|
||||||
|
|
||||||
- **Type**: boolean
|
- **Type**: boolean
|
||||||
@@ -78,7 +89,7 @@ Project Configuration Values
|
|||||||
|
|
||||||
- **Type**: list of strings
|
- **Type**: list of strings
|
||||||
- **Default**: ``[]``
|
- **Default**: ``[]``
|
||||||
- **Description**: A list of pages to ignore.
|
- **Description**: A list of pages to ignore using glob patterns.
|
||||||
See :ref:`excluding_content`.
|
See :ref:`excluding_content`.
|
||||||
|
|
||||||
.. versionadded:: 0.2.1
|
.. versionadded:: 0.2.1
|
||||||
|
|||||||
@@ -1,11 +1,6 @@
|
|||||||
Getting Started
|
Getting Started
|
||||||
===============
|
===============
|
||||||
|
|
||||||
Demo
|
|
||||||
----
|
|
||||||
|
|
||||||
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
|
||||||
|
|
||||||
Installation
|
Installation
|
||||||
------------
|
------------
|
||||||
|
|
||||||
@@ -15,6 +10,12 @@ Directly install via ``pip`` by using:
|
|||||||
|
|
||||||
pip install sphinx-llms-txt
|
pip install sphinx-llms-txt
|
||||||
|
|
||||||
|
Or with ``conda`` via ``conda-forge``:
|
||||||
|
|
||||||
|
.. code::
|
||||||
|
|
||||||
|
conda install -c conda-forge sphinx-llms-txt
|
||||||
|
|
||||||
Usage
|
Usage
|
||||||
-----
|
-----
|
||||||
|
|
||||||
@@ -26,25 +27,12 @@ Add the extension to your Sphinx configuration (``conf.py``):
|
|||||||
'sphinx_llms_txt',
|
'sphinx_llms_txt',
|
||||||
]
|
]
|
||||||
|
|
||||||
Once added, the extension will automatically generate the LLMs.txt files during the build process.
|
After the HTML finishes building, **sphinx-llms-txt** will output the location of the output files::
|
||||||
|
|
||||||
|
sphinx-llms-txt: Created /path/to/_build/html/llms-full.txt with 45 sources and 6879 lines
|
||||||
|
sphinx-llms-txt: created /path/to/_build/html/llms.txt
|
||||||
|
|
||||||
|
|
||||||
|
.. tip:: Make sure to confirm the accuracy of the output files after installs and upgrades.
|
||||||
|
|
||||||
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
|
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
|
||||||
|
|
||||||
How It Works
|
|
||||||
------------
|
|
||||||
|
|
||||||
During the Sphinx build process:
|
|
||||||
|
|
||||||
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
|
|
||||||
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
|
|
||||||
3. **Path Resolution**: Transforms relative paths in directives to full paths
|
|
||||||
4. **Output Generation**: Creates two optional files:
|
|
||||||
|
|
||||||
- ``llms.txt``: A concise summary of your documentation, in Markdown
|
|
||||||
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
|
|
||||||
|
|
||||||
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
|
|
||||||
|
|
||||||
|
|
||||||
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
|
||||||
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
|
||||||
|
|||||||
@@ -5,6 +5,25 @@ A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Mar
|
|||||||
|
|
||||||
|PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars|
|
|PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars|
|
||||||
|
|
||||||
|
Demo
|
||||||
|
----
|
||||||
|
|
||||||
|
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
|
||||||
|
|
||||||
|
Highlights
|
||||||
|
----------
|
||||||
|
|
||||||
|
1. **Content Collection**: Quickly gathers content from _sources, without needing a separate build
|
||||||
|
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
|
||||||
|
3. **Path Resolution**: Transforms relative paths in directives to full paths
|
||||||
|
4. **Output Generation**: Creates two optional files:
|
||||||
|
|
||||||
|
- ``llms.txt``: A concise summary of your documentation, in Markdown
|
||||||
|
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
|
||||||
|
|
||||||
|
5. **Content Filtering**: Allows you to exclude specific pages or sections
|
||||||
|
6. **Source Code**: Allows you to include specific source code files
|
||||||
|
|
||||||
.. toctree::
|
.. toctree::
|
||||||
:maxdepth: 2
|
:maxdepth: 2
|
||||||
|
|
||||||
@@ -15,6 +34,8 @@ A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Mar
|
|||||||
changelog
|
changelog
|
||||||
|
|
||||||
|
|
||||||
|
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
|
||||||
|
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
|
||||||
.. _Sphinx: http://sphinx-doc.org/
|
.. _Sphinx: http://sphinx-doc.org/
|
||||||
|
|
||||||
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
|
||||||
|
|||||||
@@ -1,5 +1,14 @@
|
|||||||
"""
|
"""
|
||||||
Sphinx extension to create a combined sources file (llms-full.txt)
|
Sphinx extension that generates llms.txt and llms-full.txt files for LLM consumption.
|
||||||
|
|
||||||
|
This extension collects documentation content from Sphinx projects and generates
|
||||||
|
two output files:
|
||||||
|
- llms.txt: A concise Markdown summary with project overview and page links
|
||||||
|
- llms-full.txt: A comprehensive reStructuredText file containing all documentation
|
||||||
|
content with resolved includes and path references
|
||||||
|
|
||||||
|
The extension processes content during the build phase, handles page-level and
|
||||||
|
block-level ignore directives, and can optionally include source code files.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from typing import Any, Dict
|
from typing import Any, Dict
|
||||||
@@ -12,7 +21,7 @@ from .manager import LLMSFullManager
|
|||||||
from .processor import DocumentProcessor
|
from .processor import DocumentProcessor
|
||||||
from .writer import FileWriter
|
from .writer import FileWriter
|
||||||
|
|
||||||
__version__ = "0.4.0"
|
__version__ = "0.5.0"
|
||||||
|
|
||||||
# Export classes needed by tests
|
# Export classes needed by tests
|
||||||
__all__ = [
|
__all__ = [
|
||||||
@@ -33,6 +42,13 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
|
|||||||
"""Called when a docname has been resolved to a document."""
|
"""Called when a docname has been resolved to a document."""
|
||||||
global _root_first_paragraph
|
global _root_first_paragraph
|
||||||
|
|
||||||
|
# Check for llms-txt-ignore metadata at the page level
|
||||||
|
if hasattr(app.env, "metadata") and docname in app.env.metadata:
|
||||||
|
metadata = app.env.metadata[docname]
|
||||||
|
if metadata.get("llms-txt-ignore", "").lower() in ("true", "1", "yes"):
|
||||||
|
_manager.mark_page_ignored(docname)
|
||||||
|
return
|
||||||
|
|
||||||
# Extract title from the document
|
# Extract title from the document
|
||||||
title = None
|
title = None
|
||||||
# findall() returns a generator, convert to list to check if it has elements
|
# findall() returns a generator, convert to list to check if it has elements
|
||||||
@@ -74,6 +90,7 @@ def build_finished(app: Sphinx, exception):
|
|||||||
"llms_txt_full_file": app.config.llms_txt_full_file,
|
"llms_txt_full_file": app.config.llms_txt_full_file,
|
||||||
"llms_txt_full_filename": app.config.llms_txt_full_filename,
|
"llms_txt_full_filename": app.config.llms_txt_full_filename,
|
||||||
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
"llms_txt_full_max_size": app.config.llms_txt_full_max_size,
|
||||||
|
"llms_txt_full_size_policy": app.config.llms_txt_full_size_policy,
|
||||||
"llms_txt_directives": app.config.llms_txt_directives,
|
"llms_txt_directives": app.config.llms_txt_directives,
|
||||||
"llms_txt_exclude": app.config.llms_txt_exclude,
|
"llms_txt_exclude": app.config.llms_txt_exclude,
|
||||||
"llms_txt_code_files": app.config.llms_txt_code_files,
|
"llms_txt_code_files": app.config.llms_txt_code_files,
|
||||||
@@ -90,7 +107,7 @@ def build_finished(app: Sphinx, exception):
|
|||||||
_manager.update_page_title(docname, title)
|
_manager.update_page_title(docname, title)
|
||||||
|
|
||||||
# Create the combined file
|
# Create the combined file
|
||||||
_manager.combine_sources(app.outdir, app.srcdir)
|
_manager.combine_sources(str(app.outdir), str(app.srcdir))
|
||||||
|
|
||||||
|
|
||||||
def setup(app: Sphinx) -> Dict[str, Any]:
|
def setup(app: Sphinx) -> Dict[str, Any]:
|
||||||
@@ -102,6 +119,7 @@ def setup(app: Sphinx) -> Dict[str, Any]:
|
|||||||
app.add_config_value("llms_txt_full_file", True, "env")
|
app.add_config_value("llms_txt_full_file", True, "env")
|
||||||
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
|
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
|
||||||
app.add_config_value("llms_txt_full_max_size", None, "env")
|
app.add_config_value("llms_txt_full_max_size", None, "env")
|
||||||
|
app.add_config_value("llms_txt_full_size_policy", "warn_skip", "env")
|
||||||
app.add_config_value("llms_txt_directives", [], "env")
|
app.add_config_value("llms_txt_directives", [], "env")
|
||||||
app.add_config_value("llms_txt_title", None, "env")
|
app.add_config_value("llms_txt_title", None, "env")
|
||||||
app.add_config_value("llms_txt_summary", None, "env")
|
app.add_config_value("llms_txt_summary", None, "env")
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
import fnmatch
|
import fnmatch
|
||||||
from typing import Any, Dict, List, Tuple
|
from typing import Any, Dict, List, Optional, Tuple
|
||||||
|
|
||||||
from sphinx.environment import BuildEnvironment
|
from sphinx.environment import BuildEnvironment
|
||||||
from sphinx.util import logging
|
from sphinx.util import logging
|
||||||
@@ -16,8 +16,8 @@ class DocumentCollector:
|
|||||||
|
|
||||||
def __init__(self):
|
def __init__(self):
|
||||||
self.page_titles: Dict[str, str] = {}
|
self.page_titles: Dict[str, str] = {}
|
||||||
self.master_doc: str = None
|
self.master_doc: Optional[str] = None
|
||||||
self.env: BuildEnvironment = None
|
self.env: Optional[BuildEnvironment] = None
|
||||||
self.config: Dict[str, Any] = {}
|
self.config: Dict[str, Any] = {}
|
||||||
self.app = None
|
self.app = None
|
||||||
|
|
||||||
@@ -60,7 +60,7 @@ class DocumentCollector:
|
|||||||
else:
|
else:
|
||||||
return [source_suffix] # String format
|
return [source_suffix] # String format
|
||||||
|
|
||||||
def _get_docname_suffix(self, docname: str, sources_dir) -> str:
|
def _get_docname_suffix(self, docname: str, sources_dir) -> Optional[str]:
|
||||||
"""
|
"""
|
||||||
Determine the source suffix for a given docname by checking which
|
Determine the source suffix for a given docname by checking which
|
||||||
file exists.
|
file exists.
|
||||||
@@ -102,7 +102,7 @@ class DocumentCollector:
|
|||||||
|
|
||||||
return None
|
return None
|
||||||
|
|
||||||
def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
|
def get_page_order(self, sources_dir=None) -> List[Tuple[str, Optional[str]]]:
|
||||||
"""Get the correct page order from the toctree structure.
|
"""Get the correct page order from the toctree structure.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
@@ -114,7 +114,7 @@ class DocumentCollector:
|
|||||||
if not self.env or not self.master_doc:
|
if not self.env or not self.master_doc:
|
||||||
return []
|
return []
|
||||||
|
|
||||||
page_order = []
|
page_order: List[Tuple[str, Optional[str]]] = []
|
||||||
visited = set()
|
visited = set()
|
||||||
|
|
||||||
def collect_from_toctree(docname: str):
|
def collect_from_toctree(docname: str):
|
||||||
@@ -126,7 +126,7 @@ class DocumentCollector:
|
|||||||
|
|
||||||
# Add the current document with its suffix
|
# Add the current document with its suffix
|
||||||
if docname not in [doc for doc, _ in page_order]:
|
if docname not in [doc for doc, _ in page_order]:
|
||||||
suffix = None
|
suffix: Optional[str] = None
|
||||||
if sources_dir:
|
if sources_dir:
|
||||||
suffix = self._get_docname_suffix(docname, sources_dir)
|
suffix = self._get_docname_suffix(docname, sources_dir)
|
||||||
page_order.append((docname, suffix))
|
page_order.append((docname, suffix))
|
||||||
@@ -135,18 +135,21 @@ class DocumentCollector:
|
|||||||
try:
|
try:
|
||||||
# Look for toctree_includes which contains the direct children
|
# Look for toctree_includes which contains the direct children
|
||||||
if (
|
if (
|
||||||
hasattr(self.env, "toctree_includes")
|
self.env
|
||||||
|
and hasattr(self.env, "toctree_includes")
|
||||||
and docname in self.env.toctree_includes
|
and docname in self.env.toctree_includes
|
||||||
):
|
):
|
||||||
for child_docname in self.env.toctree_includes[docname]:
|
for child_docname in self.env.toctree_includes[docname]:
|
||||||
collect_from_toctree(child_docname)
|
collect_from_toctree(str(child_docname))
|
||||||
# Try to use dependencies to find related documents
|
# Try to use dependencies to find related documents
|
||||||
elif (
|
elif (
|
||||||
hasattr(self.env, "dependencies")
|
self.env
|
||||||
|
and hasattr(self.env, "dependencies")
|
||||||
and docname in self.env.dependencies
|
and docname in self.env.dependencies
|
||||||
):
|
):
|
||||||
# Extract the dependent documents from the dependencies dict
|
# Extract the dependent documents from the dependencies dict
|
||||||
for child_docname in self.env.dependencies[docname]:
|
for child_docname_obj in self.env.dependencies[docname]:
|
||||||
|
child_docname = str(child_docname_obj)
|
||||||
# Only add documents actually in the document set
|
# Only add documents actually in the document set
|
||||||
if (
|
if (
|
||||||
hasattr(self.env, "all_docs")
|
hasattr(self.env, "all_docs")
|
||||||
@@ -154,7 +157,11 @@ class DocumentCollector:
|
|||||||
):
|
):
|
||||||
collect_from_toctree(child_docname)
|
collect_from_toctree(child_docname)
|
||||||
# Fallback to titles or other available references
|
# Fallback to titles or other available references
|
||||||
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
|
elif (
|
||||||
|
self.env
|
||||||
|
and hasattr(self.env, "titles")
|
||||||
|
and hasattr(self.env, "all_docs")
|
||||||
|
):
|
||||||
# Get all document names
|
# Get all document names
|
||||||
all_docnames = list(self.env.all_docs.keys())
|
all_docnames = list(self.env.all_docs.keys())
|
||||||
|
|
||||||
@@ -185,7 +192,7 @@ class DocumentCollector:
|
|||||||
]
|
]
|
||||||
)
|
)
|
||||||
for docname in remaining:
|
for docname in remaining:
|
||||||
suffix = None
|
suffix: Optional[str] = None
|
||||||
if sources_dir:
|
if sources_dir:
|
||||||
suffix = self._get_docname_suffix(docname, sources_dir)
|
suffix = self._get_docname_suffix(docname, sources_dir)
|
||||||
page_order.append((docname, suffix))
|
page_order.append((docname, suffix))
|
||||||
@@ -193,8 +200,8 @@ class DocumentCollector:
|
|||||||
return page_order
|
return page_order
|
||||||
|
|
||||||
def filter_excluded_pages(
|
def filter_excluded_pages(
|
||||||
self, page_order: List[Tuple[str, str]]
|
self, page_order: List[Tuple[str, Optional[str]]]
|
||||||
) -> List[Tuple[str, str]]:
|
) -> List[Tuple[str, Optional[str]]]:
|
||||||
"""Filter out excluded pages from the page order."""
|
"""Filter out excluded pages from the page order."""
|
||||||
exclude_patterns = self.config.get("llms_txt_exclude")
|
exclude_patterns = self.config.get("llms_txt_exclude")
|
||||||
if exclude_patterns:
|
if exclude_patterns:
|
||||||
|
|||||||
+198
-30
@@ -5,7 +5,7 @@ Main manager module for sphinx-llms-txt.
|
|||||||
import glob
|
import glob
|
||||||
import subprocess
|
import subprocess
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Dict, List, Optional, Tuple
|
from typing import Any, Dict, List, Optional, Tuple, Union, cast
|
||||||
|
|
||||||
from sphinx.application import Sphinx
|
from sphinx.application import Sphinx
|
||||||
from sphinx.environment import BuildEnvironment
|
from sphinx.environment import BuildEnvironment
|
||||||
@@ -129,6 +129,7 @@ class LLMSFullManager:
|
|||||||
self.srcdir: Optional[str] = None
|
self.srcdir: Optional[str] = None
|
||||||
self.outdir: Optional[str] = None
|
self.outdir: Optional[str] = None
|
||||||
self.app: Optional[Sphinx] = None
|
self.app: Optional[Sphinx] = None
|
||||||
|
self.ignored_pages: set = set()
|
||||||
|
|
||||||
def set_master_doc(self, master_doc: str):
|
def set_master_doc(self, master_doc: str):
|
||||||
"""Set the master document name."""
|
"""Set the master document name."""
|
||||||
@@ -144,6 +145,27 @@ class LLMSFullManager:
|
|||||||
"""Update the title for a page."""
|
"""Update the title for a page."""
|
||||||
self.collector.update_page_title(docname, title)
|
self.collector.update_page_title(docname, title)
|
||||||
|
|
||||||
|
def mark_page_ignored(self, docname: str):
|
||||||
|
"""Mark a page as ignored due to llms-txt-ignore metadata."""
|
||||||
|
self.ignored_pages.add(docname)
|
||||||
|
|
||||||
|
def _filter_ignored_pages(
|
||||||
|
self, page_order: Union[List[str], List[Tuple[str, Optional[str]]]]
|
||||||
|
) -> Union[List[str], List[Tuple[str, Optional[str]]]]:
|
||||||
|
"""Filter out ignored pages from page_order."""
|
||||||
|
filtered_pages = []
|
||||||
|
for item in page_order:
|
||||||
|
# Handle both old format (str) and new format (tuple)
|
||||||
|
if isinstance(item, tuple):
|
||||||
|
docname, _ = item
|
||||||
|
else:
|
||||||
|
docname = item
|
||||||
|
|
||||||
|
if docname not in self.ignored_pages:
|
||||||
|
filtered_pages.append(item)
|
||||||
|
|
||||||
|
return cast(Union[List[str], List[Tuple[str, Optional[str]]]], filtered_pages)
|
||||||
|
|
||||||
def set_config(self, config: Dict[str, Any]):
|
def set_config(self, config: Dict[str, Any]):
|
||||||
"""Set configuration options."""
|
"""Set configuration options."""
|
||||||
self.config = config
|
self.config = config
|
||||||
@@ -203,7 +225,7 @@ class LLMSFullManager:
|
|||||||
|
|
||||||
# Determine output file name and location
|
# Determine output file name and location
|
||||||
output_filename = self.config.get("llms_txt_full_filename")
|
output_filename = self.config.get("llms_txt_full_filename")
|
||||||
output_path = Path(outdir) / output_filename
|
output_path = Path(outdir) / str(output_filename)
|
||||||
|
|
||||||
# Log discovered files and page order
|
# Log discovered files and page order
|
||||||
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
|
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
|
||||||
@@ -264,7 +286,7 @@ class LLMSFullManager:
|
|||||||
content_parts = []
|
content_parts = []
|
||||||
|
|
||||||
# Track code files for later processing
|
# Track code files for later processing
|
||||||
code_file_parts = []
|
code_file_parts: List[str] = []
|
||||||
|
|
||||||
# Count lines in code files (initially 0)
|
# Count lines in code files (initially 0)
|
||||||
code_files_line_count = 0
|
code_files_line_count = 0
|
||||||
@@ -273,16 +295,39 @@ class LLMSFullManager:
|
|||||||
added_files = set()
|
added_files = set()
|
||||||
total_line_count = code_files_line_count
|
total_line_count = code_files_line_count
|
||||||
max_lines = self.config.get("llms_txt_full_max_size")
|
max_lines = self.config.get("llms_txt_full_max_size")
|
||||||
abort_due_to_max_lines = False
|
|
||||||
|
# Parse size_policy configuration early to determine collection strategy
|
||||||
|
size_policy_action = None
|
||||||
|
aborted_due_to_size = False
|
||||||
|
if max_lines is not None:
|
||||||
|
size_policy = self.config.get("llms_txt_full_size_policy", "warn_skip")
|
||||||
|
_, size_policy_action = self._parse_size_policy_config(size_policy)
|
||||||
|
|
||||||
|
# Only collect all files if action is "keep"
|
||||||
|
# For "skip" and "note", we can abort early when size limit is exceeded
|
||||||
|
should_abort_early = size_policy_action in ["skip", "note"]
|
||||||
|
|
||||||
for docname, _ in page_order:
|
for docname, _ in page_order:
|
||||||
|
# Skip pages marked as ignored
|
||||||
|
if docname in self.ignored_pages:
|
||||||
|
logger.debug(f"sphinx-llms-txt: Skipping ignored page: {docname}")
|
||||||
|
continue
|
||||||
|
|
||||||
if docname in docname_to_file:
|
if docname in docname_to_file:
|
||||||
file_path = docname_to_file[docname]
|
file_path = docname_to_file[docname]
|
||||||
content, line_count = self._read_source_file(file_path, docname)
|
content, line_count = self._read_source_file(file_path, docname)
|
||||||
|
|
||||||
# Check if adding this file would exceed the maximum line count
|
# Abort early for skip/note actions
|
||||||
if max_lines is not None and total_line_count + line_count > max_lines:
|
if (
|
||||||
abort_due_to_max_lines = True
|
max_lines is not None
|
||||||
|
and total_line_count + line_count > max_lines
|
||||||
|
and should_abort_early
|
||||||
|
):
|
||||||
|
logger.debug(
|
||||||
|
f"sphinx-llms-txt: Stopping collection due to size limit. "
|
||||||
|
f"File {docname} would exceed limit."
|
||||||
|
)
|
||||||
|
aborted_due_to_size = True
|
||||||
break
|
break
|
||||||
|
|
||||||
# Double-check this file should be included (not in excluded patterns)
|
# Double-check this file should be included (not in excluded patterns)
|
||||||
@@ -315,10 +360,12 @@ class LLMSFullManager:
|
|||||||
)
|
)
|
||||||
|
|
||||||
# Add any remaining files (in alphabetical order) that aren't in the page order
|
# Add any remaining files (in alphabetical order) that aren't in the page order
|
||||||
if not abort_due_to_max_lines:
|
# Only skip this if we aborted early due to size limits for skip/note actions
|
||||||
|
size_limit_exceeded = max_lines is not None and total_line_count > max_lines
|
||||||
|
if not (size_limit_exceeded and should_abort_early):
|
||||||
# Get all source files in the _sources directory using configured suffixes
|
# Get all source files in the _sources directory using configured suffixes
|
||||||
source_suffixes = self._get_source_suffixes()
|
source_suffixes = self._get_source_suffixes()
|
||||||
all_source_files = []
|
all_source_files: List[Path] = []
|
||||||
for src_suffix in source_suffixes:
|
for src_suffix in source_suffixes:
|
||||||
# Avoid duplicate extensions when source_suffix == source_link_suffix
|
# Avoid duplicate extensions when source_suffix == source_link_suffix
|
||||||
if src_suffix == source_link_suffix:
|
if src_suffix == source_link_suffix:
|
||||||
@@ -363,6 +410,13 @@ class LLMSFullManager:
|
|||||||
if docname is None:
|
if docname is None:
|
||||||
continue
|
continue
|
||||||
|
|
||||||
|
# Skip pages marked as ignored
|
||||||
|
if docname in self.ignored_pages:
|
||||||
|
logger.debug(
|
||||||
|
f"sphinx-llms-txt: Skipping ignored remaining file: {docname}"
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
|
||||||
# Skip excluded docnames
|
# Skip excluded docnames
|
||||||
if exclude_patterns and any(
|
if exclude_patterns and any(
|
||||||
self.collector._match_exclude_pattern(docname, pattern)
|
self.collector._match_exclude_pattern(docname, pattern)
|
||||||
@@ -374,8 +428,13 @@ class LLMSFullManager:
|
|||||||
# Read and process the file
|
# Read and process the file
|
||||||
content, line_count = self._read_source_file(file_path, docname)
|
content, line_count = self._read_source_file(file_path, docname)
|
||||||
|
|
||||||
# Check if adding this file would exceed the maximum line count
|
# Abort early for skip/note actions
|
||||||
if max_lines is not None and total_line_count + line_count > max_lines:
|
if (
|
||||||
|
max_lines is not None
|
||||||
|
and total_line_count + line_count > max_lines
|
||||||
|
and should_abort_early
|
||||||
|
):
|
||||||
|
aborted_due_to_size = True
|
||||||
break
|
break
|
||||||
|
|
||||||
if content:
|
if content:
|
||||||
@@ -384,23 +443,26 @@ class LLMSFullManager:
|
|||||||
total_line_count += line_count
|
total_line_count += line_count
|
||||||
|
|
||||||
# Process code files at the end if configured
|
# Process code files at the end if configured
|
||||||
if not abort_due_to_max_lines:
|
# Only skip this if we aborted early due to size limits for skip/note actions
|
||||||
|
if not (size_limit_exceeded and should_abort_early):
|
||||||
code_file_parts, processed_file_paths = self._process_code_files()
|
code_file_parts, processed_file_paths = self._process_code_files()
|
||||||
code_files_line_count = sum(
|
code_files_line_count = sum(
|
||||||
part.count("\n") + 1 for part in code_file_parts
|
part.count("\n") + 1 for part in code_file_parts
|
||||||
)
|
)
|
||||||
|
|
||||||
# Check if adding code files would exceed the maximum line count
|
# Check if adding code files would exceed the maximum line count
|
||||||
max_lines = self.config.get("llms_txt_full_max_size")
|
# For "keep" action, we include code files regardless of size
|
||||||
if (
|
if (
|
||||||
max_lines is not None
|
max_lines is not None
|
||||||
and total_line_count + code_files_line_count > max_lines
|
and total_line_count + code_files_line_count > max_lines
|
||||||
|
and should_abort_early
|
||||||
):
|
):
|
||||||
logger.warning(
|
logger.warning(
|
||||||
f"sphinx-llms-txt: Adding code files would exceed max line limit "
|
f"sphinx-llms-txt: Adding code files would exceed max line limit "
|
||||||
f"({max_lines}). Current: {total_line_count}, "
|
f"({max_lines}). Current: {total_line_count}, "
|
||||||
f"Code files: {code_files_line_count}. Skipping code files."
|
f"Code files: {code_files_line_count}. Skipping code files."
|
||||||
)
|
)
|
||||||
|
aborted_due_to_size = True
|
||||||
else:
|
else:
|
||||||
# Add source code files section if there are any code files
|
# Add source code files section if there are any code files
|
||||||
if code_file_parts:
|
if code_file_parts:
|
||||||
@@ -413,35 +475,70 @@ class LLMSFullManager:
|
|||||||
total_line_count += (
|
total_line_count += (
|
||||||
code_files_line_count + section_header.count("\n") + 1
|
code_files_line_count + section_header.count("\n") + 1
|
||||||
)
|
)
|
||||||
|
else:
|
||||||
|
# If we aborted early for skip/note actions, set empty code file parts
|
||||||
|
code_file_parts = []
|
||||||
|
|
||||||
# Check if line limit was exceeded before creating the file
|
# Handle size limit exceeded cases
|
||||||
max_lines = self.config.get("llms_txt_full_max_size")
|
if max_lines is not None and (
|
||||||
if abort_due_to_max_lines or (
|
total_line_count > max_lines or aborted_due_to_size
|
||||||
max_lines is not None and total_line_count > max_lines
|
|
||||||
):
|
):
|
||||||
logger.warning(
|
# Parse the size_policy configuration (reuse what we parsed earlier)
|
||||||
f"sphinx-llms-txt: Max line limit ({max_lines}) exceeded:"
|
size_policy = self.config.get("llms_txt_full_size_policy", "warn_skip")
|
||||||
f" {total_line_count} > {max_lines}. "
|
log_level, action = self._parse_size_policy_config(size_policy)
|
||||||
f"Not creating llms-full.txt file."
|
|
||||||
|
# Log with the specified level
|
||||||
|
filename = self.config.get("llms_txt_full_filename", "llms-full.txt")
|
||||||
|
message = f"sphinx-llms-txt: Max lines ({max_lines}) exceeded for {filename}" # noqa: E501
|
||||||
|
|
||||||
|
if log_level == "info":
|
||||||
|
logger.info(message)
|
||||||
|
else:
|
||||||
|
logger.warning(message)
|
||||||
|
|
||||||
|
# Handle different actions
|
||||||
|
if action == "skip":
|
||||||
|
filename = self.config.get("llms_txt_full_filename", "llms-full.txt")
|
||||||
|
logger.info(f"sphinx-llms-txt: Skipping {filename} generation")
|
||||||
|
# Log summary information if requested
|
||||||
|
if self.config.get("llms_txt_file"):
|
||||||
|
filtered_page_order = self._filter_ignored_pages(page_order)
|
||||||
|
self.writer.write_verbose_info_to_file(
|
||||||
|
filtered_page_order,
|
||||||
|
self.collector.page_titles,
|
||||||
|
total_line_count,
|
||||||
)
|
)
|
||||||
|
return
|
||||||
|
elif action == "note":
|
||||||
|
logger.info(f"sphinx-llms-txt: Creating placeholder {output_path}")
|
||||||
|
self._write_placeholder_file(output_path, max_lines)
|
||||||
|
|
||||||
# Log summary information if requested
|
# Log summary information if requested
|
||||||
if self.config.get("llms_txt_file"):
|
if self.config.get("llms_txt_file"):
|
||||||
|
filtered_page_order = self._filter_ignored_pages(page_order)
|
||||||
self.writer.write_verbose_info_to_file(
|
self.writer.write_verbose_info_to_file(
|
||||||
page_order, self.collector.page_titles, total_line_count
|
filtered_page_order,
|
||||||
|
self.collector.page_titles,
|
||||||
|
total_line_count,
|
||||||
)
|
)
|
||||||
|
|
||||||
return
|
return
|
||||||
|
elif action == "keep":
|
||||||
|
filename = self.config.get("llms_txt_full_filename", "llms-full.txt")
|
||||||
|
# Fall through to write the file
|
||||||
|
|
||||||
# Write combined file if limit wasn't exceeded
|
# Write combined file only if we have content to write
|
||||||
|
if content_parts:
|
||||||
success = self.writer.write_combined_file(
|
success = self.writer.write_combined_file(
|
||||||
content_parts, output_path, total_line_count
|
content_parts, output_path, total_line_count
|
||||||
)
|
)
|
||||||
|
else:
|
||||||
|
success = False
|
||||||
|
|
||||||
# Log summary information if requested
|
# Log summary information if requested
|
||||||
if success and self.config.get("llms_txt_file"):
|
if success and self.config.get("llms_txt_file"):
|
||||||
|
filtered_page_order = self._filter_ignored_pages(page_order)
|
||||||
self.writer.write_verbose_info_to_file(
|
self.writer.write_verbose_info_to_file(
|
||||||
page_order, self.collector.page_titles, total_line_count
|
filtered_page_order, self.collector.page_titles, total_line_count
|
||||||
)
|
)
|
||||||
|
|
||||||
def _read_source_file(self, file_path: Path, docname: str) -> Tuple[str, int]:
|
def _read_source_file(self, file_path: Path, docname: str) -> Tuple[str, int]:
|
||||||
@@ -638,9 +735,9 @@ class LLMSFullManager:
|
|||||||
title = Path(title_str[len(base_path) :])
|
title = Path(title_str[len(base_path) :])
|
||||||
except ValueError:
|
except ValueError:
|
||||||
# File is not relative to srcdir, use filename
|
# File is not relative to srcdir, use filename
|
||||||
title = file_path.name
|
title = Path(file_path.name)
|
||||||
else:
|
else:
|
||||||
title = file_path.name
|
title = Path(file_path.name)
|
||||||
|
|
||||||
# Format as code block with equals underline
|
# Format as code block with equals underline
|
||||||
title_str = str(title)
|
title_str = str(title)
|
||||||
@@ -672,7 +769,9 @@ class LLMSFullManager:
|
|||||||
|
|
||||||
return code_parts, sorted(processed_files)
|
return code_parts, sorted(processed_files)
|
||||||
|
|
||||||
def _create_code_files_section_header(self, file_paths: List[Path] = None) -> str:
|
def _create_code_files_section_header(
|
||||||
|
self, file_paths: Optional[List[Path]] = None
|
||||||
|
) -> str:
|
||||||
"""Create the section header for source code files.
|
"""Create the section header for source code files.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
@@ -720,7 +819,7 @@ class LLMSFullManager:
|
|||||||
return ""
|
return ""
|
||||||
|
|
||||||
# Convert to relative paths if possible and create tree structure
|
# Convert to relative paths if possible and create tree structure
|
||||||
tree_data = {}
|
tree_data: Dict[str, Any] = {}
|
||||||
|
|
||||||
for file_path in sorted(file_paths):
|
for file_path in sorted(file_paths):
|
||||||
# Get relative path from source directory for display
|
# Get relative path from source directory for display
|
||||||
@@ -773,7 +872,7 @@ class LLMSFullManager:
|
|||||||
current[parts[-1]] = None # None indicates it's a file
|
current[parts[-1]] = None # None indicates it's a file
|
||||||
|
|
||||||
# Convert tree structure to string representation
|
# Convert tree structure to string representation
|
||||||
lines = []
|
lines: List[str] = []
|
||||||
self._format_tree_node(tree_data, lines, "", True)
|
self._format_tree_node(tree_data, lines, "", True)
|
||||||
|
|
||||||
# Indent each line for reStructuredText code block
|
# Indent each line for reStructuredText code block
|
||||||
@@ -813,3 +912,72 @@ class LLMSFullManager:
|
|||||||
# Recursively handle subdirectories
|
# Recursively handle subdirectories
|
||||||
if subtree is not None: # It's a directory
|
if subtree is not None: # It's a directory
|
||||||
self._format_tree_node(subtree, lines, next_prefix, False)
|
self._format_tree_node(subtree, lines, next_prefix, False)
|
||||||
|
|
||||||
|
def _parse_size_policy_config(self, size_policy: str) -> tuple[str, str]:
|
||||||
|
"""Parse the llms_txt_full_size_policy configuration value.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
size_policy: Configuration string in format "loglevel_action"
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (log_level, action) where:
|
||||||
|
- log_level is "warn" or "info"
|
||||||
|
- action is "keep", "skip", or "note"
|
||||||
|
"""
|
||||||
|
if not size_policy or "_" not in size_policy:
|
||||||
|
logger.warning(
|
||||||
|
f"sphinx-llms-txt: Invalid llms_txt_full_size_policy "
|
||||||
|
f"format: '{size_policy}'. "
|
||||||
|
f"Using default 'warn_skip'."
|
||||||
|
)
|
||||||
|
return "warn", "skip"
|
||||||
|
|
||||||
|
parts = size_policy.split("_", 1) # Split on first underscore only
|
||||||
|
log_level, action = parts[0], parts[1]
|
||||||
|
|
||||||
|
# Validate log level
|
||||||
|
if log_level not in ["warn", "info"]:
|
||||||
|
logger.warning(
|
||||||
|
f"sphinx-llms-txt: Invalid log level '{log_level}' in "
|
||||||
|
f"llms_txt_full_size_policy. "
|
||||||
|
f"Valid options: warn, info. Using 'warn'."
|
||||||
|
)
|
||||||
|
log_level = "warn"
|
||||||
|
|
||||||
|
# Validate action
|
||||||
|
if action not in ["keep", "skip", "note"]:
|
||||||
|
logger.warning(
|
||||||
|
f"sphinx-llms-txt: Invalid action '{action}' in "
|
||||||
|
f"llms_txt_full_size_policy. "
|
||||||
|
f"Valid options: keep, skip, note. Using 'skip'."
|
||||||
|
)
|
||||||
|
action = "skip"
|
||||||
|
|
||||||
|
return log_level, action
|
||||||
|
|
||||||
|
def _write_placeholder_file(self, output_path: Path, max_lines: int):
|
||||||
|
"""Write a placeholder llms-full.txt file with a note about size limit.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
output_path: Path where the placeholder file should be written
|
||||||
|
max_lines: The configured maximum line limit
|
||||||
|
"""
|
||||||
|
# Create the placeholder note content
|
||||||
|
placeholder_content = (
|
||||||
|
f".. This file was not generated because it exceeded the configured size limit.\n" # noqa: E501
|
||||||
|
" See the conf.py ``llms_txt_full_max_size`` and ``llms_txt_full_size_policy``\n" # noqa: E501
|
||||||
|
" for configuration options.\n"
|
||||||
|
"\n"
|
||||||
|
f" Configured max size: {max_lines} lines\n"
|
||||||
|
"\n"
|
||||||
|
" For more information, see: https://sphinx-llms-txt.readthedocs.io/en/latest/configuration-values.html#llms-txt-full-max-size\n" # noqa: E501
|
||||||
|
)
|
||||||
|
|
||||||
|
try:
|
||||||
|
with open(output_path, "w", encoding="utf-8") as f:
|
||||||
|
f.write(placeholder_content)
|
||||||
|
logger.debug(f"sphinx-llms-txt: Wrote placeholder file: {output_path}")
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(
|
||||||
|
f"sphinx-llms-txt: Error writing placeholder file {output_path}: {e}"
|
||||||
|
)
|
||||||
|
|||||||
@@ -44,7 +44,10 @@ class DocumentProcessor:
|
|||||||
Returns:
|
Returns:
|
||||||
Processed content with directives properly resolved
|
Processed content with directives properly resolved
|
||||||
"""
|
"""
|
||||||
# First process include directives
|
# First process llms-txt-ignore blocks
|
||||||
|
content = self._process_ignore_blocks(content)
|
||||||
|
|
||||||
|
# Then process include directives
|
||||||
content = self._process_includes(content, source_path)
|
content = self._process_includes(content, source_path)
|
||||||
|
|
||||||
# Then process path directives (image, figure, etc.)
|
# Then process path directives (image, figure, etc.)
|
||||||
@@ -127,7 +130,7 @@ class DocumentProcessor:
|
|||||||
"""
|
"""
|
||||||
# Get the configured path directives to process
|
# Get the configured path directives to process
|
||||||
default_path_directives = ["image", "figure"]
|
default_path_directives = ["image", "figure"]
|
||||||
custom_path_directives = self.config.get("llms_txt_directives")
|
custom_path_directives = self.config.get("llms_txt_directives") or []
|
||||||
path_directives = set(default_path_directives + custom_path_directives)
|
path_directives = set(default_path_directives + custom_path_directives)
|
||||||
|
|
||||||
# Build the regex pattern to match all configured directives
|
# Build the regex pattern to match all configured directives
|
||||||
@@ -243,9 +246,12 @@ class DocumentProcessor:
|
|||||||
"""
|
"""
|
||||||
possible_paths = []
|
possible_paths = []
|
||||||
|
|
||||||
# If it's an absolute path, use it directly
|
# If it's an absolute path, treat it as relative to srcdir
|
||||||
if os.path.isabs(include_path):
|
if os.path.isabs(include_path):
|
||||||
possible_paths.append(Path(include_path))
|
# Remove the leading slash and treat as relative to srcdir
|
||||||
|
relative_path = include_path.lstrip("/")
|
||||||
|
if self.srcdir:
|
||||||
|
possible_paths.append((Path(self.srcdir) / relative_path).resolve())
|
||||||
else:
|
else:
|
||||||
# Relative to the source file (in _sources directory)
|
# Relative to the source file (in _sources directory)
|
||||||
possible_paths.append((source_path.parent / include_path).resolve())
|
possible_paths.append((source_path.parent / include_path).resolve())
|
||||||
@@ -286,6 +292,9 @@ class DocumentProcessor:
|
|||||||
# Function to replace each include with content
|
# Function to replace each include with content
|
||||||
def replace_include(match):
|
def replace_include(match):
|
||||||
include_path = match.group(3)
|
include_path = match.group(3)
|
||||||
|
directive_part = match.group(
|
||||||
|
1
|
||||||
|
) # The ".. include:: " part with leading whitespace
|
||||||
|
|
||||||
# Get all possible paths to try
|
# Get all possible paths to try
|
||||||
possible_paths = self._resolve_include_paths(include_path, source_path)
|
possible_paths = self._resolve_include_paths(include_path, source_path)
|
||||||
@@ -296,7 +305,18 @@ class DocumentProcessor:
|
|||||||
if path_to_try.exists():
|
if path_to_try.exists():
|
||||||
with open(path_to_try, "r", encoding="utf-8") as f:
|
with open(path_to_try, "r", encoding="utf-8") as f:
|
||||||
included_content = f.read()
|
included_content = f.read()
|
||||||
|
|
||||||
|
# Find where the actual directive starts, after any whitespace
|
||||||
|
directive_start = directive_part.find("..")
|
||||||
|
if directive_start > 0:
|
||||||
|
# There's leading whitespace/newlines before the directive
|
||||||
|
leading_part = directive_part[:directive_start]
|
||||||
|
# Replace directive with content, preserving the structure
|
||||||
|
return leading_part + included_content
|
||||||
|
else:
|
||||||
|
# No leading whitespace, just return the content
|
||||||
return included_content
|
return included_content
|
||||||
|
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.error(
|
logger.error(
|
||||||
f"sphinx-llms-txt: Error reading include file {path_to_try}:"
|
f"sphinx-llms-txt: Error reading include file {path_to_try}:"
|
||||||
@@ -308,8 +328,48 @@ class DocumentProcessor:
|
|||||||
paths_tried = ", ".join(str(p) for p in possible_paths)
|
paths_tried = ", ".join(str(p) for p in possible_paths)
|
||||||
logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}")
|
logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}")
|
||||||
logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}")
|
logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}")
|
||||||
|
|
||||||
|
# Preserve spacing structure for error message too
|
||||||
|
directive_start = match.group(1).find("..")
|
||||||
|
if directive_start > 0:
|
||||||
|
leading_part = match.group(1)[:directive_start]
|
||||||
|
return leading_part + f"[Include file not found: {include_path}]"
|
||||||
|
else:
|
||||||
return f"[Include file not found: {include_path}]"
|
return f"[Include file not found: {include_path}]"
|
||||||
|
|
||||||
# Replace all includes with their content
|
# Replace all includes with their content
|
||||||
processed_content = include_pattern.sub(replace_include, content)
|
processed_content = include_pattern.sub(replace_include, content)
|
||||||
return processed_content
|
return processed_content
|
||||||
|
|
||||||
|
def _process_ignore_blocks(self, content: str) -> str:
|
||||||
|
"""Process llms-txt-ignore-start/end blocks by removing their content.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
content: The source content to process
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Processed content with ignore blocks removed
|
||||||
|
"""
|
||||||
|
# Process ignore blocks iteratively to handle nested cases correctly
|
||||||
|
while True:
|
||||||
|
# Pattern to match ignore blocks - handles whitespace and indentation
|
||||||
|
ignore_pattern = re.compile(
|
||||||
|
r"^\s*\.\.\s+llms-txt-ignore-start\s*\n" # Start directive line
|
||||||
|
r"(.*?)" # Content to ignore (non-greedy)
|
||||||
|
r"^\s*\.\.\s+llms-txt-ignore-end\s*$", # End directive line
|
||||||
|
re.MULTILINE | re.DOTALL,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Find and remove one ignore block at a time
|
||||||
|
match = ignore_pattern.search(content)
|
||||||
|
if not match:
|
||||||
|
break
|
||||||
|
|
||||||
|
# Remove the matched block
|
||||||
|
content = content[: match.start()] + content[match.end() :]
|
||||||
|
|
||||||
|
# Clean up any extra blank lines that might be left
|
||||||
|
# Replace multiple consecutive newlines with at most 2 newlines
|
||||||
|
processed_content = re.sub(r"\n\n\n+", "\n\n", content)
|
||||||
|
|
||||||
|
return processed_content
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Dict, List, Tuple, Union
|
from typing import Any, Dict, List, Optional, Tuple, Union
|
||||||
|
|
||||||
from sphinx.application import Sphinx
|
from sphinx.application import Sphinx
|
||||||
from sphinx.util import logging
|
from sphinx.util import logging
|
||||||
@@ -14,7 +14,12 @@ logger = logging.getLogger(__name__)
|
|||||||
class FileWriter:
|
class FileWriter:
|
||||||
"""Handles writing processed content to output files."""
|
"""Handles writing processed content to output files."""
|
||||||
|
|
||||||
def __init__(self, config: Dict[str, Any], outdir: str = None, app: Sphinx = None):
|
def __init__(
|
||||||
|
self,
|
||||||
|
config: Dict[str, Any],
|
||||||
|
outdir: Optional[str] = None,
|
||||||
|
app: Optional[Sphinx] = None,
|
||||||
|
):
|
||||||
self.config = config
|
self.config = config
|
||||||
self.outdir = outdir
|
self.outdir = outdir
|
||||||
self.app = app
|
self.app = app
|
||||||
@@ -37,7 +42,7 @@ class FileWriter:
|
|||||||
f.write("\n".join(content_parts))
|
f.write("\n".join(content_parts))
|
||||||
|
|
||||||
logger.info(
|
logger.info(
|
||||||
f"sphinx-llms-txt: created {output_path} with {len(content_parts)}"
|
f"sphinx-llms-txt: Created {output_path} with {len(content_parts)}"
|
||||||
f" sources and {total_line_count} lines"
|
f" sources and {total_line_count} lines"
|
||||||
)
|
)
|
||||||
return True
|
return True
|
||||||
@@ -47,7 +52,7 @@ class FileWriter:
|
|||||||
|
|
||||||
def write_verbose_info_to_file(
|
def write_verbose_info_to_file(
|
||||||
self,
|
self,
|
||||||
page_order: Union[List[str], List[Tuple[str, str]]],
|
page_order: Union[List[str], List[Tuple[str, Optional[str]]]],
|
||||||
page_titles: Dict[str, str],
|
page_titles: Dict[str, str],
|
||||||
total_line_count: int = 0,
|
total_line_count: int = 0,
|
||||||
) -> bool:
|
) -> bool:
|
||||||
@@ -67,13 +72,13 @@ class FileWriter:
|
|||||||
)
|
)
|
||||||
return False
|
return False
|
||||||
|
|
||||||
output_path = Path(self.outdir) / self.config.get("llms_txt_filename")
|
output_path = Path(self.outdir) / str(self.config.get("llms_txt_filename"))
|
||||||
try:
|
try:
|
||||||
with open(output_path, "w", encoding="utf-8") as f:
|
with open(output_path, "w", encoding="utf-8") as f:
|
||||||
project_name = "llms-txt Summary"
|
project_name = "llms-txt Summary"
|
||||||
# First priority: use title from config if available
|
# First priority: use title from config if available
|
||||||
if self.config.get("llms_txt_title"):
|
if self.config.get("llms_txt_title"):
|
||||||
project_name = self.config.get("llms_txt_title")
|
project_name = str(self.config.get("llms_txt_title"))
|
||||||
# Second priority: use project name from Sphinx app if available
|
# Second priority: use project name from Sphinx app if available
|
||||||
elif (
|
elif (
|
||||||
self.app
|
self.app
|
||||||
|
|||||||
@@ -8,6 +8,8 @@ Welcome to Test Project's documentation!
|
|||||||
page1
|
page1
|
||||||
page2
|
page2
|
||||||
page_with_include
|
page_with_include
|
||||||
|
page_ignored_metadata
|
||||||
|
page_with_ignore_blocks
|
||||||
|
|
||||||
Indices and tables
|
Indices and tables
|
||||||
==================
|
==================
|
||||||
|
|||||||
@@ -0,0 +1,16 @@
|
|||||||
|
:llms-txt-ignore: true
|
||||||
|
|
||||||
|
Page Ignored by Metadata
|
||||||
|
========================
|
||||||
|
|
||||||
|
This page should not appear in llms-full.txt because of the metadata directive.
|
||||||
|
|
||||||
|
Section 1
|
||||||
|
---------
|
||||||
|
|
||||||
|
This content should be completely ignored.
|
||||||
|
|
||||||
|
Section 2
|
||||||
|
---------
|
||||||
|
|
||||||
|
This content should also be ignored.
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
Page With Ignore Blocks
|
||||||
|
=======================
|
||||||
|
|
||||||
|
This content should appear in llms-full.txt.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
This content should be ignored and not appear in llms-full.txt.
|
||||||
|
|
||||||
|
Section Ignored
|
||||||
|
---------------
|
||||||
|
|
||||||
|
This section should also be ignored.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
This content after the ignore block should appear in llms-full.txt.
|
||||||
|
|
||||||
|
Another Section
|
||||||
|
---------------
|
||||||
|
|
||||||
|
This content should definitely appear.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
Another ignored block with multiple lines.
|
||||||
|
|
||||||
|
- Item 1 (ignored)
|
||||||
|
- Item 2 (ignored)
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
# This code should be ignored
|
||||||
|
def ignored_function():
|
||||||
|
pass
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
Final content that should appear.
|
||||||
@@ -0,0 +1,220 @@
|
|||||||
|
"""Tests for llms-txt ignore features."""
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from sphinx_llms_txt import DocumentProcessor
|
||||||
|
|
||||||
|
|
||||||
|
def test_process_ignore_blocks():
|
||||||
|
"""Test that ignore blocks are properly removed from content."""
|
||||||
|
processor = DocumentProcessor({}, None)
|
||||||
|
|
||||||
|
content = """This content should remain.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
This content should be removed.
|
||||||
|
|
||||||
|
Section Ignored
|
||||||
|
---------------
|
||||||
|
|
||||||
|
This section should also be removed.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
This content should remain after the ignore block.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
Another ignored block.
|
||||||
|
Multiple lines here.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
Final content that should remain."""
|
||||||
|
|
||||||
|
processed = processor._process_ignore_blocks(content)
|
||||||
|
|
||||||
|
# Check that ignored content is removed
|
||||||
|
assert "This content should be removed." not in processed
|
||||||
|
assert "Section Ignored" not in processed
|
||||||
|
assert "Another ignored block." not in processed
|
||||||
|
assert "Multiple lines here." not in processed
|
||||||
|
|
||||||
|
# Check that non-ignored content remains
|
||||||
|
assert "This content should remain." in processed
|
||||||
|
assert "This content should remain after the ignore block." in processed
|
||||||
|
assert "Final content that should remain." in processed
|
||||||
|
|
||||||
|
|
||||||
|
def test_process_ignore_blocks_with_indentation():
|
||||||
|
"""Test that ignore blocks work with different indentation levels."""
|
||||||
|
processor = DocumentProcessor({}, None)
|
||||||
|
|
||||||
|
content = """Section Title
|
||||||
|
=============
|
||||||
|
|
||||||
|
Normal content.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
Indented ignored content.
|
||||||
|
More indented content.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
Back to normal content."""
|
||||||
|
|
||||||
|
processed = processor._process_ignore_blocks(content)
|
||||||
|
|
||||||
|
# Check that ignored content is removed
|
||||||
|
assert "Indented ignored content." not in processed
|
||||||
|
assert "More indented content." not in processed
|
||||||
|
|
||||||
|
# Check that non-ignored content remains
|
||||||
|
assert "Section Title" in processed
|
||||||
|
assert "Normal content." in processed
|
||||||
|
assert "Back to normal content." in processed
|
||||||
|
|
||||||
|
|
||||||
|
def test_process_ignore_blocks_multiple():
|
||||||
|
"""Test that multiple ignore blocks are handled correctly."""
|
||||||
|
processor = DocumentProcessor({}, None)
|
||||||
|
|
||||||
|
content = """Start content.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
First ignore block.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
Middle content that should remain.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
Second ignore block.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
End content."""
|
||||||
|
|
||||||
|
processed = processor._process_ignore_blocks(content)
|
||||||
|
|
||||||
|
# Check that ignored content is removed
|
||||||
|
assert "First ignore block." not in processed
|
||||||
|
assert "Second ignore block." not in processed
|
||||||
|
|
||||||
|
# Check that non-ignored content remains
|
||||||
|
assert "Start content." in processed
|
||||||
|
assert "Middle content that should remain." in processed
|
||||||
|
assert "End content." in processed
|
||||||
|
|
||||||
|
|
||||||
|
def test_build_with_ignore_features(basic_sphinx_app):
|
||||||
|
"""Test building HTML documentation with ignore features."""
|
||||||
|
app = basic_sphinx_app
|
||||||
|
app.build()
|
||||||
|
|
||||||
|
# Check if the output file was created
|
||||||
|
output_file = Path(app.outdir) / "test-llms-full.txt"
|
||||||
|
assert output_file.exists(), f"Output file {output_file} does not exist"
|
||||||
|
|
||||||
|
# Read the content of the output file
|
||||||
|
content = output_file.read_text()
|
||||||
|
|
||||||
|
# Check that page with metadata ignore is completely excluded
|
||||||
|
assert "Page Ignored by Metadata" not in content
|
||||||
|
assert "This page should not appear in llms-full.txt" not in content
|
||||||
|
|
||||||
|
# Check that page with ignore blocks has the right content
|
||||||
|
assert "Page With Ignore Blocks" in content
|
||||||
|
assert "This content should appear in llms-full.txt." in content
|
||||||
|
assert "This content after the ignore block should appear" in content
|
||||||
|
assert "Another Section" in content
|
||||||
|
assert "Final content that should appear." in content
|
||||||
|
|
||||||
|
# Check that ignored block content is not present
|
||||||
|
assert "This content should be ignored and not appear" not in content
|
||||||
|
assert "Section Ignored" not in content
|
||||||
|
assert "Another ignored block with multiple lines." not in content
|
||||||
|
assert "Item 1 (ignored)" not in content
|
||||||
|
assert "def ignored_function():" not in content
|
||||||
|
|
||||||
|
|
||||||
|
def test_manager_mark_page_ignored():
|
||||||
|
"""Test that manager can mark pages as ignored."""
|
||||||
|
from sphinx_llms_txt import LLMSFullManager
|
||||||
|
|
||||||
|
manager = LLMSFullManager()
|
||||||
|
|
||||||
|
# Initially no pages are ignored
|
||||||
|
assert len(manager.ignored_pages) == 0
|
||||||
|
|
||||||
|
# Mark a page as ignored
|
||||||
|
manager.mark_page_ignored("test_page")
|
||||||
|
|
||||||
|
# Check that page is in ignored set
|
||||||
|
assert "test_page" in manager.ignored_pages
|
||||||
|
assert len(manager.ignored_pages) == 1
|
||||||
|
|
||||||
|
# Mark another page as ignored
|
||||||
|
manager.mark_page_ignored("another_page")
|
||||||
|
|
||||||
|
# Check both pages are ignored
|
||||||
|
assert "test_page" in manager.ignored_pages
|
||||||
|
assert "another_page" in manager.ignored_pages
|
||||||
|
assert len(manager.ignored_pages) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_process_ignore_blocks_empty_blocks():
|
||||||
|
"""Test that empty ignore blocks are handled correctly."""
|
||||||
|
processor = DocumentProcessor({}, None)
|
||||||
|
|
||||||
|
content = """Content before.
|
||||||
|
|
||||||
|
.. llms-txt-ignore-start
|
||||||
|
|
||||||
|
.. llms-txt-ignore-end
|
||||||
|
|
||||||
|
Content after."""
|
||||||
|
|
||||||
|
processed = processor._process_ignore_blocks(content)
|
||||||
|
|
||||||
|
# Check that content remains
|
||||||
|
assert "Content before." in processed
|
||||||
|
assert "Content after." in processed
|
||||||
|
|
||||||
|
# Check that we don't have excessive newlines
|
||||||
|
lines = processed.strip().split("\n")
|
||||||
|
non_empty_lines = [line for line in lines if line.strip()]
|
||||||
|
assert len(non_empty_lines) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_ignore_metadata_affects_both_files(basic_sphinx_app):
|
||||||
|
"""Test that :llms-txt-ignore: true affects both files."""
|
||||||
|
app = basic_sphinx_app
|
||||||
|
# Enable both llms.txt and llms-full.txt file generation
|
||||||
|
app.config.llms_txt_file = True
|
||||||
|
app.config.llms_txt_filename = "test-llms.txt"
|
||||||
|
app.build()
|
||||||
|
|
||||||
|
# Check if both output files were created
|
||||||
|
llms_full_file = Path(app.outdir) / "test-llms-full.txt"
|
||||||
|
llms_summary_file = Path(app.outdir) / "test-llms.txt"
|
||||||
|
|
||||||
|
assert llms_full_file.exists(), f"Output file {llms_full_file} does not exist"
|
||||||
|
assert llms_summary_file.exists(), f"Output file {llms_summary_file} does not exist"
|
||||||
|
|
||||||
|
# Read the content of both files
|
||||||
|
llms_full_content = llms_full_file.read_text()
|
||||||
|
llms_summary_content = llms_summary_file.read_text()
|
||||||
|
|
||||||
|
# Check that page with metadata ignore is excluded from llms-full.txt
|
||||||
|
assert "Page Ignored by Metadata" not in llms_full_content
|
||||||
|
assert "This page should not appear in llms-full.txt" not in llms_full_content
|
||||||
|
|
||||||
|
# Check that page with metadata ignore is also excluded from llms.txt
|
||||||
|
# This should NOT contain a link to the ignored page
|
||||||
|
assert "Page Ignored by Metadata" not in llms_summary_content
|
||||||
|
assert "page_ignored_metadata.html" not in llms_summary_content
|
||||||
@@ -117,6 +117,152 @@ def test_max_lines_limit(temp_dir, rootdir):
|
|||||||
app.docutils_conf_path.unlink()
|
app.docutils_conf_path.unlink()
|
||||||
|
|
||||||
|
|
||||||
|
def test_on_exceed_skip(temp_dir, rootdir):
|
||||||
|
"""Test that skip action works when size limit is exceeded."""
|
||||||
|
from sphinx.testing.util import SphinxTestApp
|
||||||
|
|
||||||
|
src_dir = rootdir / "basic"
|
||||||
|
|
||||||
|
app = SphinxTestApp(
|
||||||
|
srcdir=src_dir,
|
||||||
|
builddir=temp_dir,
|
||||||
|
buildername="html",
|
||||||
|
freshenv=True,
|
||||||
|
confoverrides={
|
||||||
|
"llms_txt_full_filename": "skip-test.txt",
|
||||||
|
"llms_txt_full_max_size": 20,
|
||||||
|
"llms_txt_full_size_policy": "warn_skip",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
app.build()
|
||||||
|
|
||||||
|
# Check that the output file was NOT created
|
||||||
|
output_file = Path(app.outdir) / "skip-test.txt"
|
||||||
|
assert (
|
||||||
|
not output_file.exists()
|
||||||
|
), f"Output file {output_file} should not exist with skip action"
|
||||||
|
|
||||||
|
# Cleanup
|
||||||
|
sys.path[:] = app._saved_path
|
||||||
|
_clean_up_global_state()
|
||||||
|
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
||||||
|
app.docutils_conf_path.unlink()
|
||||||
|
|
||||||
|
|
||||||
|
def test_on_exceed_keep(temp_dir, rootdir):
|
||||||
|
"""Test that keep action works when size limit is exceeded."""
|
||||||
|
from sphinx.testing.util import SphinxTestApp
|
||||||
|
|
||||||
|
src_dir = rootdir / "basic"
|
||||||
|
|
||||||
|
app = SphinxTestApp(
|
||||||
|
srcdir=src_dir,
|
||||||
|
builddir=temp_dir,
|
||||||
|
buildername="html",
|
||||||
|
freshenv=True,
|
||||||
|
confoverrides={
|
||||||
|
"llms_txt_full_filename": "keep-test.txt",
|
||||||
|
"llms_txt_full_max_size": 20,
|
||||||
|
"llms_txt_full_size_policy": "info_keep",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
app.build()
|
||||||
|
|
||||||
|
# Check that the output file WAS created despite exceeding limit
|
||||||
|
output_file = Path(app.outdir) / "keep-test.txt"
|
||||||
|
assert (
|
||||||
|
output_file.exists()
|
||||||
|
), f"Output file {output_file} should exist with keep action"
|
||||||
|
|
||||||
|
# Verify it has content
|
||||||
|
content = output_file.read_text()
|
||||||
|
assert len(content) > 0, "Output file should have content with keep action"
|
||||||
|
|
||||||
|
# Cleanup
|
||||||
|
sys.path[:] = app._saved_path
|
||||||
|
_clean_up_global_state()
|
||||||
|
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
||||||
|
app.docutils_conf_path.unlink()
|
||||||
|
|
||||||
|
|
||||||
|
def test_on_exceed_note(temp_dir, rootdir):
|
||||||
|
"""Test that note action works when size limit is exceeded."""
|
||||||
|
from sphinx.testing.util import SphinxTestApp
|
||||||
|
|
||||||
|
src_dir = rootdir / "basic"
|
||||||
|
|
||||||
|
app = SphinxTestApp(
|
||||||
|
srcdir=src_dir,
|
||||||
|
builddir=temp_dir,
|
||||||
|
buildername="html",
|
||||||
|
freshenv=True,
|
||||||
|
confoverrides={
|
||||||
|
"llms_txt_full_filename": "note-test.txt",
|
||||||
|
"llms_txt_full_max_size": 20,
|
||||||
|
"llms_txt_full_size_policy": "warn_note",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
app.build()
|
||||||
|
|
||||||
|
# Check that the output file WAS created with placeholder content
|
||||||
|
output_file = Path(app.outdir) / "note-test.txt"
|
||||||
|
assert (
|
||||||
|
output_file.exists()
|
||||||
|
), f"Output file {output_file} should exist with note action"
|
||||||
|
|
||||||
|
# Verify it has the placeholder content
|
||||||
|
content = output_file.read_text()
|
||||||
|
assert (
|
||||||
|
"This file was not generated because it exceeded the configured size limit."
|
||||||
|
in content
|
||||||
|
)
|
||||||
|
assert "llms_txt_full_max_size" in content
|
||||||
|
assert "llms_txt_full_size_policy" in content
|
||||||
|
assert "Configured max size: 20 lines" in content
|
||||||
|
|
||||||
|
# Cleanup
|
||||||
|
sys.path[:] = app._saved_path
|
||||||
|
_clean_up_global_state()
|
||||||
|
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
||||||
|
app.docutils_conf_path.unlink()
|
||||||
|
|
||||||
|
|
||||||
|
def test_on_exceed_invalid_config(temp_dir, rootdir):
|
||||||
|
"""Test behavior with invalid configuration values."""
|
||||||
|
from sphinx.testing.util import SphinxTestApp
|
||||||
|
|
||||||
|
src_dir = rootdir / "basic"
|
||||||
|
|
||||||
|
app = SphinxTestApp(
|
||||||
|
srcdir=src_dir,
|
||||||
|
builddir=temp_dir,
|
||||||
|
buildername="html",
|
||||||
|
freshenv=True,
|
||||||
|
confoverrides={
|
||||||
|
"llms_txt_full_filename": "invalid-test.txt",
|
||||||
|
"llms_txt_full_max_size": 20,
|
||||||
|
"llms_txt_full_size_policy": "invalid_format", # Invalid config
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
app.build()
|
||||||
|
|
||||||
|
# Should fall back to default behavior (warn_skip)
|
||||||
|
output_file = Path(app.outdir) / "invalid-test.txt"
|
||||||
|
assert (
|
||||||
|
not output_file.exists()
|
||||||
|
), f"Output file {output_file} should not exist with invalid config fallback"
|
||||||
|
|
||||||
|
# Cleanup
|
||||||
|
sys.path[:] = app._saved_path
|
||||||
|
_clean_up_global_state()
|
||||||
|
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
|
||||||
|
app.docutils_conf_path.unlink()
|
||||||
|
|
||||||
|
|
||||||
def test_title_override(temp_dir, rootdir):
|
def test_title_override(temp_dir, rootdir):
|
||||||
"""Test that the title override works correctly."""
|
"""Test that the title override works correctly."""
|
||||||
from sphinx.testing.util import SphinxTestApp
|
from sphinx.testing.util import SphinxTestApp
|
||||||
|
|||||||
@@ -762,6 +762,7 @@ def test_summary_default_uses_first_paragraph():
|
|||||||
llms_txt_full_file = True
|
llms_txt_full_file = True
|
||||||
llms_txt_full_filename = "llms-full.txt"
|
llms_txt_full_filename = "llms-full.txt"
|
||||||
llms_txt_full_max_size = None
|
llms_txt_full_max_size = None
|
||||||
|
llms_txt_full_size_policy = "warn_skip"
|
||||||
llms_txt_directives = []
|
llms_txt_directives = []
|
||||||
llms_txt_exclude = []
|
llms_txt_exclude = []
|
||||||
llms_txt_code_files = []
|
llms_txt_code_files = []
|
||||||
|
|||||||
Reference in New Issue
Block a user