Compare commits

..
10 Commits
19 changed files with 930 additions and 84 deletions
+1 -1
View File
@@ -33,7 +33,7 @@ jobs:
- name: Install dependencies - name: Install dependencies
run: | run: |
python -m pip install --upgrade pip python -m pip install --upgrade pip
pip install -e ".[dev]" pip install -e . --group dev
# - name: Run mypy # - name: Run mypy
# run: | # run: |
+32
View File
@@ -1,6 +1,38 @@
Changelog Changelog
========= =========
0.5.3
-----
- Make sphinx a required dependency since there are imports from Sphinx
`#44 <https://github.com/jdillard/sphinx-llms-txt/pull/44>`_
0.5.2
-----
- Remove support for singlehtml
`#40 <https://github.com/jdillard/sphinx-llms-txt/pull/40>`_
0.5.1
-----
- Only allow builders that have _sources directory
`#38 <https://github.com/jdillard/sphinx-llms-txt/pull/38>`_
0.5.0
-----
- Add :ref:`block_level_ignore` and :ref:`page_level_ignore`
`#33 <https://github.com/jdillard/sphinx-llms-txt/pull/33>`_
- Add :confval:`llms_txt_full_size_policy` configuration option to control behavior when :confval:`llms_txt_full_max_size` is exceeded.
`#35 <https://github.com/jdillard/sphinx-llms-txt/pull/35>`_
0.4.1
-----
- Fix include paths and spacing
`#31 <https://github.com/jdillard/sphinx-llms-txt/pull/31>`_
0.4.0 0.4.0
----- -----
+4
View File
@@ -11,6 +11,10 @@ A Sphinx extension that generates a summary `llms.txt` file and a single combine
See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions. See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
## Contributing
Pull Requests welcome! See [Contributing](https://sphinx-llms-txt.readthedocs.io/en/latest/contributing.html) for instructions on how best to contribute.
## License ## License
MIT License - see LICENSE file for details. MIT License - see LICENSE file for details.
+83 -3
View File
@@ -76,15 +76,26 @@ Handling Large Documentation
^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
For very large documentation sets, generating the full documentation file might exceed reasonable size limits. For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
You can set a maximum line count: You can set a maximum line count and control what happens when that limit is exceeded:
.. code-block:: python .. code-block:: python
llms_txt_full_max_size = 10000 # Maximum 10,000 lines llms_txt_full_max_size = 10000 # Maximum 10,000 lines
llms_txt_full_size_policy = "warn_skip" # Default behavior
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete. The ``llms_txt_full_size_policy`` setting controls both the log level and action taken when the size limit is exceeded.
It uses the format ``"<loglevel>_<action>"``:
.. tip:: Use :ref:`excluding_content` to remove less relevant pages. **Log levels:**
- ``warn``: Log as a warning (default)
- ``info``: Log as informational message
**Actions:**
- ``skip``: Don't create the file (default)
- ``keep``: Create the file anyway, ignoring the size limit
- ``note``: Create a placeholder file explaining why the full file wasn't generated
.. tip:: Use :ref:`excluding_content` to remove less relevant pages and reduce the file size.
.. _custom_directive_handling: .. _custom_directive_handling:
@@ -113,6 +124,13 @@ This ensures that paths in your custom directives are properly resolved in the g
Excluding Content Excluding Content
^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^^
There are several ways to exclude content from the generated ``llms-full.txt`` file:
.. _global_exclusion:
Global Page Exclusion
~~~~~~~~~~~~~~~~~~~~~~
You can exclude specific pages from being included in the generated files: You can exclude specific pages from being included in the generated files:
.. code-block:: python .. code-block:: python
@@ -124,6 +142,67 @@ You can exclude specific pages from being included in the generated files:
] ]
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption. This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
It can also be used to reduce the size of llms-full.txt.
.. _page_level_ignore:
Page-Level Ignore Metadata
~~~~~~~~~~~~~~~~~~~~~~~~~~~
You can exclude individual pages by adding metadata at the top of any reStructuredText file:
.. code-block:: restructuredtext
:llms-txt-ignore: true
Page Title
==========
This entire page will be excluded from llms-full.txt
When this metadata is present, the entire page is skipped during processing.
.. _block_level_ignore:
Block-Level Ignore Directives
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
You can exclude specific sections within a page using ignore directives:
.. code-block:: restructuredtext
Page Title
==========
This content will be included in llms-full.txt.
.. llms-txt-ignore-start
This content will be excluded from llms-full.txt.
Section To Ignore
-----------------
This entire section and any nested content will be ignored.
.. code-block:: python
# This code block will also be ignored
def ignored_function():
pass
.. llms-txt-ignore-end
This content will be included again.
Block-level ignores can be useful for:
- Removing internal notes or TODOs
- Hiding implementation details while keeping user-facing documentation
.. note::
- Multiple ignore blocks can be used within the same file
- Ignore directives work with any indentation level
.. _including_code_files: .. _including_code_files:
@@ -198,6 +277,7 @@ Here's a complete example showing multiple :doc:`configuration-values`:
llms_txt_filename = "ai-summary.txt" llms_txt_filename = "ai-summary.txt"
llms_txt_full_filename = "ai-full-docs.txt" llms_txt_full_filename = "ai-full-docs.txt"
llms_txt_full_max_size = 50000 llms_txt_full_max_size = 50000
llms_txt_full_size_policy = "warn_note"
# Content customization # Content customization
llms_txt_title = "Project Documentation for AI Assistants" llms_txt_title = "Project Documentation for AI Assistants"
+13 -2
View File
@@ -24,11 +24,22 @@ Project Configuration Values
- **Type**: integer or ``None`` - **Type**: integer or ``None``
- **Default**: ``None`` (no limit) - **Default**: ``None`` (no limit)
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``. - **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
If exceeded, the file is skipped and a warning is shown, but the build still completes. Behavior when exceeded is controlled by :confval:`llms_txt_full_size_policy`.
See :ref:`handling_large_documentation`. See :ref:`handling_large_documentation`.
.. versionadded:: 0.2.0 .. versionadded:: 0.2.0
.. confval:: llms_txt_full_size_policy
- **Type**: string
- **Default**: ``'warn_skip'``
- **Description**: Controls what happens when :confval:`llms_txt_full_max_size` is exceeded.
Format is ``<loglevel>_<action>``. Log levels: ``warn``, ``info``.
Actions: ``skip``, ``keep``, ``note``.
See :ref:`handling_large_documentation`.
.. versionadded:: 0.5.0
.. confval:: llms_txt_file .. confval:: llms_txt_file
- **Type**: boolean - **Type**: boolean
@@ -78,7 +89,7 @@ Project Configuration Values
- **Type**: list of strings - **Type**: list of strings
- **Default**: ``[]`` - **Default**: ``[]``
- **Description**: A list of pages to ignore. - **Description**: A list of pages to ignore using glob patterns.
See :ref:`excluding_content`. See :ref:`excluding_content`.
.. versionadded:: 0.2.1 .. versionadded:: 0.2.1
+1 -1
View File
@@ -19,7 +19,7 @@ Local development
.. code-block:: console .. code-block:: console
pip install -e ".[dev]" pip install -e . --group dev
#. Install pre-commit Git hook scripts: #. Install pre-commit Git hook scripts:
+13 -25
View File
@@ -1,11 +1,6 @@
Getting Started Getting Started
=============== ===============
Demo
----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
Installation Installation
------------ ------------
@@ -15,6 +10,12 @@ Directly install via ``pip`` by using:
pip install sphinx-llms-txt pip install sphinx-llms-txt
Or with ``conda`` via ``conda-forge``:
.. code::
conda install -c conda-forge sphinx-llms-txt
Usage Usage
----- -----
@@ -26,25 +27,12 @@ Add the extension to your Sphinx configuration (``conf.py``):
'sphinx_llms_txt', 'sphinx_llms_txt',
] ]
Once added, the extension will automatically generate the LLMs.txt files during the build process. After the HTML finishes building, **sphinx-llms-txt** will output the location of the output files::
sphinx-llms-txt: Created /path/to/_build/html/llms-full.txt with 45 sources and 6879 lines
sphinx-llms-txt: created /path/to/_build/html/llms.txt
.. tip:: Make sure to confirm the accuracy of the output files after installs and upgrades.
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**. See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
How It Works
------------
During the Sphinx build process:
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
3. **Path Resolution**: Transforms relative paths in directives to full paths
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
+21
View File
@@ -5,6 +5,25 @@ A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Mar
|PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars| |PyPI version| |Conda Version| |Downloads| |Parallel Safe| |GitHub Stars|
Demo
----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
Highlights
----------
1. **Content Collection**: Quickly gathers content from _sources, without needing a separate build
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
3. **Path Resolution**: Transforms relative paths in directives to full paths
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
5. **Content Filtering**: Allows you to exclude specific pages or sections
6. **Source Code**: Allows you to include specific source code files
.. toctree:: .. toctree::
:maxdepth: 2 :maxdepth: 2
@@ -15,6 +34,8 @@ A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Mar
changelog changelog
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
.. _Sphinx: http://sphinx-doc.org/ .. _Sphinx: http://sphinx-doc.org/
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg .. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
+4 -2
View File
@@ -26,13 +26,16 @@ classifiers = [
license = {text = "MIT"} license = {text = "MIT"}
readme = "README.md" readme = "README.md"
dynamic = ["version"] dynamic = ["version"]
dependencies = [
"sphinx",
]
[project.urls] [project.urls]
download = "https://pypi.org/project/sphinx-llms-txt/" download = "https://pypi.org/project/sphinx-llms-txt/"
source = "https://github.com/jdillard/sphinx-llms-txt" source = "https://github.com/jdillard/sphinx-llms-txt"
changelog = "https://github.com/jdillard/sphinx-llms-txt/blob/master/CHANGELOG.rst" changelog = "https://github.com/jdillard/sphinx-llms-txt/blob/master/CHANGELOG.rst"
[project.optional-dependencies] [dependency-groups]
dev = [ dev = [
"pytest>=7.0.0", "pytest>=7.0.0",
"black", "black",
@@ -40,7 +43,6 @@ dev = [
"mypy", "mypy",
"isort", "isort",
"pre-commit", "pre-commit",
"sphinx",
] ]
test = [ test = [
"pytest>=7.0.0", "pytest>=7.0.0",
+29 -6
View File
@@ -1,5 +1,14 @@
""" """
Sphinx extension to create a combined sources file (llms-full.txt) Sphinx extension that generates llms.txt and llms-full.txt files for LLM consumption.
This extension collects documentation content from Sphinx projects and generates
two output files:
- llms.txt: A concise Markdown summary with project overview and page links
- llms-full.txt: A comprehensive reStructuredText file containing all documentation
content with resolved includes and path references
The extension processes content during the build phase, handles page-level and
block-level ignore directives, and can optionally include source code files.
""" """
from typing import Any, Dict from typing import Any, Dict
@@ -12,7 +21,7 @@ from .manager import LLMSFullManager
from .processor import DocumentProcessor from .processor import DocumentProcessor
from .writer import FileWriter from .writer import FileWriter
__version__ = "0.4.0" __version__ = "0.5.3"
# Export classes needed by tests # Export classes needed by tests
__all__ = [ __all__ = [
@@ -33,6 +42,13 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
"""Called when a docname has been resolved to a document.""" """Called when a docname has been resolved to a document."""
global _root_first_paragraph global _root_first_paragraph
# Check for llms-txt-ignore metadata at the page level
if hasattr(app.env, "metadata") and docname in app.env.metadata:
metadata = app.env.metadata[docname]
if metadata.get("llms-txt-ignore", "").lower() in ("true", "1", "yes"):
_manager.mark_page_ignored(docname)
return
# Extract title from the document # Extract title from the document
title = None title = None
# findall() returns a generator, convert to list to check if it has elements # findall() returns a generator, convert to list to check if it has elements
@@ -74,6 +90,7 @@ def build_finished(app: Sphinx, exception):
"llms_txt_full_file": app.config.llms_txt_full_file, "llms_txt_full_file": app.config.llms_txt_full_file,
"llms_txt_full_filename": app.config.llms_txt_full_filename, "llms_txt_full_filename": app.config.llms_txt_full_filename,
"llms_txt_full_max_size": app.config.llms_txt_full_max_size, "llms_txt_full_max_size": app.config.llms_txt_full_max_size,
"llms_txt_full_size_policy": app.config.llms_txt_full_size_policy,
"llms_txt_directives": app.config.llms_txt_directives, "llms_txt_directives": app.config.llms_txt_directives,
"llms_txt_exclude": app.config.llms_txt_exclude, "llms_txt_exclude": app.config.llms_txt_exclude,
"llms_txt_code_files": app.config.llms_txt_code_files, "llms_txt_code_files": app.config.llms_txt_code_files,
@@ -96,12 +113,12 @@ def build_finished(app: Sphinx, exception):
def setup(app: Sphinx) -> Dict[str, Any]: def setup(app: Sphinx) -> Dict[str, Any]:
"""Set up the Sphinx extension.""" """Set up the Sphinx extension."""
# Add configuration options
app.add_config_value("llms_txt_file", True, "env") app.add_config_value("llms_txt_file", True, "env")
app.add_config_value("llms_txt_filename", "llms.txt", "env") app.add_config_value("llms_txt_filename", "llms.txt", "env")
app.add_config_value("llms_txt_full_file", True, "env") app.add_config_value("llms_txt_full_file", True, "env")
app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env") app.add_config_value("llms_txt_full_filename", "llms-full.txt", "env")
app.add_config_value("llms_txt_full_max_size", None, "env") app.add_config_value("llms_txt_full_max_size", None, "env")
app.add_config_value("llms_txt_full_size_policy", "warn_skip", "env")
app.add_config_value("llms_txt_directives", [], "env") app.add_config_value("llms_txt_directives", [], "env")
app.add_config_value("llms_txt_title", None, "env") app.add_config_value("llms_txt_title", None, "env")
app.add_config_value("llms_txt_summary", None, "env") app.add_config_value("llms_txt_summary", None, "env")
@@ -109,15 +126,21 @@ def setup(app: Sphinx) -> Dict[str, Any]:
app.add_config_value("llms_txt_code_files", [], "env") app.add_config_value("llms_txt_code_files", [], "env")
app.add_config_value("llms_txt_code_base_path", None, "env") app.add_config_value("llms_txt_code_base_path", None, "env")
# Connect to Sphinx events def builder_inited(app):
app.connect("doctree-resolved", doctree_resolved) """Used to limit what builders are allowed to run the extension."""
app.connect("build-finished", build_finished)
allowed_builders = ["html", "dirhtml"]
if hasattr(app, "builder") and app.builder.name in allowed_builders:
# Reset manager and root paragraph for each build # Reset manager and root paragraph for each build
global _manager, _root_first_paragraph global _manager, _root_first_paragraph
_manager = LLMSFullManager() _manager = LLMSFullManager()
_root_first_paragraph = "" _root_first_paragraph = ""
app.connect("doctree-resolved", doctree_resolved)
app.connect("build-finished", build_finished)
app.connect("builder-inited", builder_inited)
return { return {
"version": __version__, "version": __version__,
"parallel_read_safe": True, "parallel_read_safe": True,
+188 -23
View File
@@ -5,7 +5,7 @@ Main manager module for sphinx-llms-txt.
import glob import glob
import subprocess import subprocess
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple from typing import Any, Dict, List, Optional, Tuple, Union
from sphinx.application import Sphinx from sphinx.application import Sphinx
from sphinx.environment import BuildEnvironment from sphinx.environment import BuildEnvironment
@@ -129,6 +129,7 @@ class LLMSFullManager:
self.srcdir: Optional[str] = None self.srcdir: Optional[str] = None
self.outdir: Optional[str] = None self.outdir: Optional[str] = None
self.app: Optional[Sphinx] = None self.app: Optional[Sphinx] = None
self.ignored_pages: set = set()
def set_master_doc(self, master_doc: str): def set_master_doc(self, master_doc: str):
"""Set the master document name.""" """Set the master document name."""
@@ -144,6 +145,27 @@ class LLMSFullManager:
"""Update the title for a page.""" """Update the title for a page."""
self.collector.update_page_title(docname, title) self.collector.update_page_title(docname, title)
def mark_page_ignored(self, docname: str):
"""Mark a page as ignored due to llms-txt-ignore metadata."""
self.ignored_pages.add(docname)
def _filter_ignored_pages(
self, page_order: Union[List[str], List[Tuple[str, str]]]
) -> Union[List[str], List[Tuple[str, str]]]:
"""Filter out ignored pages from page_order."""
filtered_pages = []
for item in page_order:
# Handle both old format (str) and new format (tuple)
if isinstance(item, tuple):
docname, _ = item
else:
docname = item
if docname not in self.ignored_pages:
filtered_pages.append(item)
return filtered_pages
def set_config(self, config: Dict[str, Any]): def set_config(self, config: Dict[str, Any]):
"""Set configuration options.""" """Set configuration options."""
self.config = config self.config = config
@@ -175,7 +197,6 @@ class LLMSFullManager:
possible_sources = [ possible_sources = [
Path(outdir) / "_sources", Path(outdir) / "_sources",
Path(outdir) / "html" / "_sources", Path(outdir) / "html" / "_sources",
Path(outdir) / "singlehtml" / "_sources",
] ]
for path in possible_sources: for path in possible_sources:
@@ -273,16 +294,39 @@ class LLMSFullManager:
added_files = set() added_files = set()
total_line_count = code_files_line_count total_line_count = code_files_line_count
max_lines = self.config.get("llms_txt_full_max_size") max_lines = self.config.get("llms_txt_full_max_size")
abort_due_to_max_lines = False
# Parse size_policy configuration early to determine collection strategy
size_policy_action = None
aborted_due_to_size = False
if max_lines is not None:
size_policy = self.config.get("llms_txt_full_size_policy", "warn_skip")
_, size_policy_action = self._parse_size_policy_config(size_policy)
# Only collect all files if action is "keep"
# For "skip" and "note", we can abort early when size limit is exceeded
should_abort_early = size_policy_action in ["skip", "note"]
for docname, _ in page_order: for docname, _ in page_order:
# Skip pages marked as ignored
if docname in self.ignored_pages:
logger.debug(f"sphinx-llms-txt: Skipping ignored page: {docname}")
continue
if docname in docname_to_file: if docname in docname_to_file:
file_path = docname_to_file[docname] file_path = docname_to_file[docname]
content, line_count = self._read_source_file(file_path, docname) content, line_count = self._read_source_file(file_path, docname)
# Check if adding this file would exceed the maximum line count # Abort early for skip/note actions
if max_lines is not None and total_line_count + line_count > max_lines: if (
abort_due_to_max_lines = True max_lines is not None
and total_line_count + line_count > max_lines
and should_abort_early
):
logger.debug(
f"sphinx-llms-txt: Stopping collection due to size limit. "
f"File {docname} would exceed limit."
)
aborted_due_to_size = True
break break
# Double-check this file should be included (not in excluded patterns) # Double-check this file should be included (not in excluded patterns)
@@ -315,7 +359,9 @@ class LLMSFullManager:
) )
# Add any remaining files (in alphabetical order) that aren't in the page order # Add any remaining files (in alphabetical order) that aren't in the page order
if not abort_due_to_max_lines: # Only skip this if we aborted early due to size limits for skip/note actions
size_limit_exceeded = max_lines is not None and total_line_count > max_lines
if not (size_limit_exceeded and should_abort_early):
# Get all source files in the _sources directory using configured suffixes # Get all source files in the _sources directory using configured suffixes
source_suffixes = self._get_source_suffixes() source_suffixes = self._get_source_suffixes()
all_source_files = [] all_source_files = []
@@ -363,6 +409,13 @@ class LLMSFullManager:
if docname is None: if docname is None:
continue continue
# Skip pages marked as ignored
if docname in self.ignored_pages:
logger.debug(
f"sphinx-llms-txt: Skipping ignored remaining file: {docname}"
)
continue
# Skip excluded docnames # Skip excluded docnames
if exclude_patterns and any( if exclude_patterns and any(
self.collector._match_exclude_pattern(docname, pattern) self.collector._match_exclude_pattern(docname, pattern)
@@ -374,8 +427,13 @@ class LLMSFullManager:
# Read and process the file # Read and process the file
content, line_count = self._read_source_file(file_path, docname) content, line_count = self._read_source_file(file_path, docname)
# Check if adding this file would exceed the maximum line count # Abort early for skip/note actions
if max_lines is not None and total_line_count + line_count > max_lines: if (
max_lines is not None
and total_line_count + line_count > max_lines
and should_abort_early
):
aborted_due_to_size = True
break break
if content: if content:
@@ -384,23 +442,26 @@ class LLMSFullManager:
total_line_count += line_count total_line_count += line_count
# Process code files at the end if configured # Process code files at the end if configured
if not abort_due_to_max_lines: # Only skip this if we aborted early due to size limits for skip/note actions
if not (size_limit_exceeded and should_abort_early):
code_file_parts, processed_file_paths = self._process_code_files() code_file_parts, processed_file_paths = self._process_code_files()
code_files_line_count = sum( code_files_line_count = sum(
part.count("\n") + 1 for part in code_file_parts part.count("\n") + 1 for part in code_file_parts
) )
# Check if adding code files would exceed the maximum line count # Check if adding code files would exceed the maximum line count
max_lines = self.config.get("llms_txt_full_max_size") # For "keep" action, we include code files regardless of size
if ( if (
max_lines is not None max_lines is not None
and total_line_count + code_files_line_count > max_lines and total_line_count + code_files_line_count > max_lines
and should_abort_early
): ):
logger.warning( logger.warning(
f"sphinx-llms-txt: Adding code files would exceed max line limit " f"sphinx-llms-txt: Adding code files would exceed max line limit "
f"({max_lines}). Current: {total_line_count}, " f"({max_lines}). Current: {total_line_count}, "
f"Code files: {code_files_line_count}. Skipping code files." f"Code files: {code_files_line_count}. Skipping code files."
) )
aborted_due_to_size = True
else: else:
# Add source code files section if there are any code files # Add source code files section if there are any code files
if code_file_parts: if code_file_parts:
@@ -413,35 +474,70 @@ class LLMSFullManager:
total_line_count += ( total_line_count += (
code_files_line_count + section_header.count("\n") + 1 code_files_line_count + section_header.count("\n") + 1
) )
else:
# If we aborted early for skip/note actions, set empty code file parts
code_file_parts = []
# Check if line limit was exceeded before creating the file # Handle size limit exceeded cases
max_lines = self.config.get("llms_txt_full_max_size") if max_lines is not None and (
if abort_due_to_max_lines or ( total_line_count > max_lines or aborted_due_to_size
max_lines is not None and total_line_count > max_lines
): ):
logger.warning( # Parse the size_policy configuration (reuse what we parsed earlier)
f"sphinx-llms-txt: Max line limit ({max_lines}) exceeded:" size_policy = self.config.get("llms_txt_full_size_policy", "warn_skip")
f" {total_line_count} > {max_lines}. " log_level, action = self._parse_size_policy_config(size_policy)
f"Not creating llms-full.txt file."
# Log with the specified level
filename = self.config.get("llms_txt_full_filename", "llms-full.txt")
message = f"sphinx-llms-txt: Max lines ({max_lines}) exceeded for {filename}" # noqa: E501
if log_level == "info":
logger.info(message)
else:
logger.warning(message)
# Handle different actions
if action == "skip":
filename = self.config.get("llms_txt_full_filename", "llms-full.txt")
logger.info(f"sphinx-llms-txt: Skipping {filename} generation")
# Log summary information if requested
if self.config.get("llms_txt_file"):
filtered_page_order = self._filter_ignored_pages(page_order)
self.writer.write_verbose_info_to_file(
filtered_page_order,
self.collector.page_titles,
total_line_count,
) )
return
elif action == "note":
logger.info(f"sphinx-llms-txt: Creating placeholder {output_path}")
self._write_placeholder_file(output_path, max_lines)
# Log summary information if requested # Log summary information if requested
if self.config.get("llms_txt_file"): if self.config.get("llms_txt_file"):
filtered_page_order = self._filter_ignored_pages(page_order)
self.writer.write_verbose_info_to_file( self.writer.write_verbose_info_to_file(
page_order, self.collector.page_titles, total_line_count filtered_page_order,
self.collector.page_titles,
total_line_count,
) )
return return
elif action == "keep":
filename = self.config.get("llms_txt_full_filename", "llms-full.txt")
# Fall through to write the file
# Write combined file if limit wasn't exceeded # Write combined file only if we have content to write
if content_parts:
success = self.writer.write_combined_file( success = self.writer.write_combined_file(
content_parts, output_path, total_line_count content_parts, output_path, total_line_count
) )
else:
success = False
# Log summary information if requested # Log summary information if requested
if success and self.config.get("llms_txt_file"): if success and self.config.get("llms_txt_file"):
filtered_page_order = self._filter_ignored_pages(page_order)
self.writer.write_verbose_info_to_file( self.writer.write_verbose_info_to_file(
page_order, self.collector.page_titles, total_line_count filtered_page_order, self.collector.page_titles, total_line_count
) )
def _read_source_file(self, file_path: Path, docname: str) -> Tuple[str, int]: def _read_source_file(self, file_path: Path, docname: str) -> Tuple[str, int]:
@@ -813,3 +909,72 @@ class LLMSFullManager:
# Recursively handle subdirectories # Recursively handle subdirectories
if subtree is not None: # It's a directory if subtree is not None: # It's a directory
self._format_tree_node(subtree, lines, next_prefix, False) self._format_tree_node(subtree, lines, next_prefix, False)
def _parse_size_policy_config(self, size_policy: str) -> tuple[str, str]:
"""Parse the llms_txt_full_size_policy configuration value.
Args:
size_policy: Configuration string in format "loglevel_action"
Returns:
Tuple of (log_level, action) where:
- log_level is "warn" or "info"
- action is "keep", "skip", or "note"
"""
if not size_policy or "_" not in size_policy:
logger.warning(
f"sphinx-llms-txt: Invalid llms_txt_full_size_policy "
f"format: '{size_policy}'. "
f"Using default 'warn_skip'."
)
return "warn", "skip"
parts = size_policy.split("_", 1) # Split on first underscore only
log_level, action = parts[0], parts[1]
# Validate log level
if log_level not in ["warn", "info"]:
logger.warning(
f"sphinx-llms-txt: Invalid log level '{log_level}' in "
f"llms_txt_full_size_policy. "
f"Valid options: warn, info. Using 'warn'."
)
log_level = "warn"
# Validate action
if action not in ["keep", "skip", "note"]:
logger.warning(
f"sphinx-llms-txt: Invalid action '{action}' in "
f"llms_txt_full_size_policy. "
f"Valid options: keep, skip, note. Using 'skip'."
)
action = "skip"
return log_level, action
def _write_placeholder_file(self, output_path: Path, max_lines: int):
"""Write a placeholder llms-full.txt file with a note about size limit.
Args:
output_path: Path where the placeholder file should be written
max_lines: The configured maximum line limit
"""
# Create the placeholder note content
placeholder_content = (
f".. This file was not generated because it exceeded the configured size limit.\n" # noqa: E501
" See the conf.py ``llms_txt_full_max_size`` and ``llms_txt_full_size_policy``\n" # noqa: E501
" for configuration options.\n"
"\n"
f" Configured max size: {max_lines} lines\n"
"\n"
" For more information, see: https://sphinx-llms-txt.readthedocs.io/en/latest/configuration-values.html#llms-txt-full-max-size\n" # noqa: E501
)
try:
with open(output_path, "w", encoding="utf-8") as f:
f.write(placeholder_content)
logger.debug(f"sphinx-llms-txt: Wrote placeholder file: {output_path}")
except Exception as e:
logger.error(
f"sphinx-llms-txt: Error writing placeholder file {output_path}: {e}"
)
+63 -3
View File
@@ -44,7 +44,10 @@ class DocumentProcessor:
Returns: Returns:
Processed content with directives properly resolved Processed content with directives properly resolved
""" """
# First process include directives # First process llms-txt-ignore blocks
content = self._process_ignore_blocks(content)
# Then process include directives
content = self._process_includes(content, source_path) content = self._process_includes(content, source_path)
# Then process path directives (image, figure, etc.) # Then process path directives (image, figure, etc.)
@@ -243,9 +246,12 @@ class DocumentProcessor:
""" """
possible_paths = [] possible_paths = []
# If it's an absolute path, use it directly # If it's an absolute path, treat it as relative to srcdir
if os.path.isabs(include_path): if os.path.isabs(include_path):
possible_paths.append(Path(include_path)) # Remove the leading slash and treat as relative to srcdir
relative_path = include_path.lstrip("/")
if self.srcdir:
possible_paths.append((Path(self.srcdir) / relative_path).resolve())
else: else:
# Relative to the source file (in _sources directory) # Relative to the source file (in _sources directory)
possible_paths.append((source_path.parent / include_path).resolve()) possible_paths.append((source_path.parent / include_path).resolve())
@@ -286,6 +292,9 @@ class DocumentProcessor:
# Function to replace each include with content # Function to replace each include with content
def replace_include(match): def replace_include(match):
include_path = match.group(3) include_path = match.group(3)
directive_part = match.group(
1
) # The ".. include:: " part with leading whitespace
# Get all possible paths to try # Get all possible paths to try
possible_paths = self._resolve_include_paths(include_path, source_path) possible_paths = self._resolve_include_paths(include_path, source_path)
@@ -296,7 +305,18 @@ class DocumentProcessor:
if path_to_try.exists(): if path_to_try.exists():
with open(path_to_try, "r", encoding="utf-8") as f: with open(path_to_try, "r", encoding="utf-8") as f:
included_content = f.read() included_content = f.read()
# Find where the actual directive starts, after any whitespace
directive_start = directive_part.find("..")
if directive_start > 0:
# There's leading whitespace/newlines before the directive
leading_part = directive_part[:directive_start]
# Replace directive with content, preserving the structure
return leading_part + included_content
else:
# No leading whitespace, just return the content
return included_content return included_content
except Exception as e: except Exception as e:
logger.error( logger.error(
f"sphinx-llms-txt: Error reading include file {path_to_try}:" f"sphinx-llms-txt: Error reading include file {path_to_try}:"
@@ -308,8 +328,48 @@ class DocumentProcessor:
paths_tried = ", ".join(str(p) for p in possible_paths) paths_tried = ", ".join(str(p) for p in possible_paths)
logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}") logger.warning(f"sphinx-llms-txt: Include file not found: {include_path}")
logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}") logger.debug(f"sphinx-llms-txt: Tried paths: {paths_tried}")
# Preserve spacing structure for error message too
directive_start = match.group(1).find("..")
if directive_start > 0:
leading_part = match.group(1)[:directive_start]
return leading_part + f"[Include file not found: {include_path}]"
else:
return f"[Include file not found: {include_path}]" return f"[Include file not found: {include_path}]"
# Replace all includes with their content # Replace all includes with their content
processed_content = include_pattern.sub(replace_include, content) processed_content = include_pattern.sub(replace_include, content)
return processed_content return processed_content
def _process_ignore_blocks(self, content: str) -> str:
"""Process llms-txt-ignore-start/end blocks by removing their content.
Args:
content: The source content to process
Returns:
Processed content with ignore blocks removed
"""
# Process ignore blocks iteratively to handle nested cases correctly
while True:
# Pattern to match ignore blocks - handles whitespace and indentation
ignore_pattern = re.compile(
r"^\s*\.\.\s+llms-txt-ignore-start\s*\n" # Start directive line
r"(.*?)" # Content to ignore (non-greedy)
r"^\s*\.\.\s+llms-txt-ignore-end\s*$", # End directive line
re.MULTILINE | re.DOTALL,
)
# Find and remove one ignore block at a time
match = ignore_pattern.search(content)
if not match:
break
# Remove the matched block
content = content[: match.start()] + content[match.end() :]
# Clean up any extra blank lines that might be left
# Replace multiple consecutive newlines with at most 2 newlines
processed_content = re.sub(r"\n\n\n+", "\n\n", content)
return processed_content
+1 -1
View File
@@ -37,7 +37,7 @@ class FileWriter:
f.write("\n".join(content_parts)) f.write("\n".join(content_parts))
logger.info( logger.info(
f"sphinx-llms-txt: created {output_path} with {len(content_parts)}" f"sphinx-llms-txt: Created {output_path} with {len(content_parts)}"
f" sources and {total_line_count} lines" f" sources and {total_line_count} lines"
) )
return True return True
+2
View File
@@ -8,6 +8,8 @@ Welcome to Test Project's documentation!
page1 page1
page2 page2
page_with_include page_with_include
page_ignored_metadata
page_with_ignore_blocks
Indices and tables Indices and tables
================== ==================
@@ -0,0 +1,16 @@
:llms-txt-ignore: true
Page Ignored by Metadata
========================
This page should not appear in llms-full.txt because of the metadata directive.
Section 1
---------
This content should be completely ignored.
Section 2
---------
This content should also be ignored.
@@ -0,0 +1,39 @@
Page With Ignore Blocks
=======================
This content should appear in llms-full.txt.
.. llms-txt-ignore-start
This content should be ignored and not appear in llms-full.txt.
Section Ignored
---------------
This section should also be ignored.
.. llms-txt-ignore-end
This content after the ignore block should appear in llms-full.txt.
Another Section
---------------
This content should definitely appear.
.. llms-txt-ignore-start
Another ignored block with multiple lines.
- Item 1 (ignored)
- Item 2 (ignored)
.. code-block:: python
# This code should be ignored
def ignored_function():
pass
.. llms-txt-ignore-end
Final content that should appear.
+220
View File
@@ -0,0 +1,220 @@
"""Tests for llms-txt ignore features."""
from pathlib import Path
from sphinx_llms_txt import DocumentProcessor
def test_process_ignore_blocks():
"""Test that ignore blocks are properly removed from content."""
processor = DocumentProcessor({}, None)
content = """This content should remain.
.. llms-txt-ignore-start
This content should be removed.
Section Ignored
---------------
This section should also be removed.
.. llms-txt-ignore-end
This content should remain after the ignore block.
.. llms-txt-ignore-start
Another ignored block.
Multiple lines here.
.. llms-txt-ignore-end
Final content that should remain."""
processed = processor._process_ignore_blocks(content)
# Check that ignored content is removed
assert "This content should be removed." not in processed
assert "Section Ignored" not in processed
assert "Another ignored block." not in processed
assert "Multiple lines here." not in processed
# Check that non-ignored content remains
assert "This content should remain." in processed
assert "This content should remain after the ignore block." in processed
assert "Final content that should remain." in processed
def test_process_ignore_blocks_with_indentation():
"""Test that ignore blocks work with different indentation levels."""
processor = DocumentProcessor({}, None)
content = """Section Title
=============
Normal content.
.. llms-txt-ignore-start
Indented ignored content.
More indented content.
.. llms-txt-ignore-end
Back to normal content."""
processed = processor._process_ignore_blocks(content)
# Check that ignored content is removed
assert "Indented ignored content." not in processed
assert "More indented content." not in processed
# Check that non-ignored content remains
assert "Section Title" in processed
assert "Normal content." in processed
assert "Back to normal content." in processed
def test_process_ignore_blocks_multiple():
"""Test that multiple ignore blocks are handled correctly."""
processor = DocumentProcessor({}, None)
content = """Start content.
.. llms-txt-ignore-start
First ignore block.
.. llms-txt-ignore-end
Middle content that should remain.
.. llms-txt-ignore-start
Second ignore block.
.. llms-txt-ignore-end
End content."""
processed = processor._process_ignore_blocks(content)
# Check that ignored content is removed
assert "First ignore block." not in processed
assert "Second ignore block." not in processed
# Check that non-ignored content remains
assert "Start content." in processed
assert "Middle content that should remain." in processed
assert "End content." in processed
def test_build_with_ignore_features(basic_sphinx_app):
"""Test building HTML documentation with ignore features."""
app = basic_sphinx_app
app.build()
# Check if the output file was created
output_file = Path(app.outdir) / "test-llms-full.txt"
assert output_file.exists(), f"Output file {output_file} does not exist"
# Read the content of the output file
content = output_file.read_text()
# Check that page with metadata ignore is completely excluded
assert "Page Ignored by Metadata" not in content
assert "This page should not appear in llms-full.txt" not in content
# Check that page with ignore blocks has the right content
assert "Page With Ignore Blocks" in content
assert "This content should appear in llms-full.txt." in content
assert "This content after the ignore block should appear" in content
assert "Another Section" in content
assert "Final content that should appear." in content
# Check that ignored block content is not present
assert "This content should be ignored and not appear" not in content
assert "Section Ignored" not in content
assert "Another ignored block with multiple lines." not in content
assert "Item 1 (ignored)" not in content
assert "def ignored_function():" not in content
def test_manager_mark_page_ignored():
"""Test that manager can mark pages as ignored."""
from sphinx_llms_txt import LLMSFullManager
manager = LLMSFullManager()
# Initially no pages are ignored
assert len(manager.ignored_pages) == 0
# Mark a page as ignored
manager.mark_page_ignored("test_page")
# Check that page is in ignored set
assert "test_page" in manager.ignored_pages
assert len(manager.ignored_pages) == 1
# Mark another page as ignored
manager.mark_page_ignored("another_page")
# Check both pages are ignored
assert "test_page" in manager.ignored_pages
assert "another_page" in manager.ignored_pages
assert len(manager.ignored_pages) == 2
def test_process_ignore_blocks_empty_blocks():
"""Test that empty ignore blocks are handled correctly."""
processor = DocumentProcessor({}, None)
content = """Content before.
.. llms-txt-ignore-start
.. llms-txt-ignore-end
Content after."""
processed = processor._process_ignore_blocks(content)
# Check that content remains
assert "Content before." in processed
assert "Content after." in processed
# Check that we don't have excessive newlines
lines = processed.strip().split("\n")
non_empty_lines = [line for line in lines if line.strip()]
assert len(non_empty_lines) == 2
def test_ignore_metadata_affects_both_files(basic_sphinx_app):
"""Test that :llms-txt-ignore: true affects both files."""
app = basic_sphinx_app
# Enable both llms.txt and llms-full.txt file generation
app.config.llms_txt_file = True
app.config.llms_txt_filename = "test-llms.txt"
app.build()
# Check if both output files were created
llms_full_file = Path(app.outdir) / "test-llms-full.txt"
llms_summary_file = Path(app.outdir) / "test-llms.txt"
assert llms_full_file.exists(), f"Output file {llms_full_file} does not exist"
assert llms_summary_file.exists(), f"Output file {llms_summary_file} does not exist"
# Read the content of both files
llms_full_content = llms_full_file.read_text()
llms_summary_content = llms_summary_file.read_text()
# Check that page with metadata ignore is excluded from llms-full.txt
assert "Page Ignored by Metadata" not in llms_full_content
assert "This page should not appear in llms-full.txt" not in llms_full_content
# Check that page with metadata ignore is also excluded from llms.txt
# This should NOT contain a link to the ignored page
assert "Page Ignored by Metadata" not in llms_summary_content
assert "page_ignored_metadata.html" not in llms_summary_content
+146
View File
@@ -117,6 +117,152 @@ def test_max_lines_limit(temp_dir, rootdir):
app.docutils_conf_path.unlink() app.docutils_conf_path.unlink()
def test_on_exceed_skip(temp_dir, rootdir):
"""Test that skip action works when size limit is exceeded."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "skip-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "warn_skip",
},
)
app.build()
# Check that the output file was NOT created
output_file = Path(app.outdir) / "skip-test.txt"
assert (
not output_file.exists()
), f"Output file {output_file} should not exist with skip action"
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_on_exceed_keep(temp_dir, rootdir):
"""Test that keep action works when size limit is exceeded."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "keep-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "info_keep",
},
)
app.build()
# Check that the output file WAS created despite exceeding limit
output_file = Path(app.outdir) / "keep-test.txt"
assert (
output_file.exists()
), f"Output file {output_file} should exist with keep action"
# Verify it has content
content = output_file.read_text()
assert len(content) > 0, "Output file should have content with keep action"
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_on_exceed_note(temp_dir, rootdir):
"""Test that note action works when size limit is exceeded."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "note-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "warn_note",
},
)
app.build()
# Check that the output file WAS created with placeholder content
output_file = Path(app.outdir) / "note-test.txt"
assert (
output_file.exists()
), f"Output file {output_file} should exist with note action"
# Verify it has the placeholder content
content = output_file.read_text()
assert (
"This file was not generated because it exceeded the configured size limit."
in content
)
assert "llms_txt_full_max_size" in content
assert "llms_txt_full_size_policy" in content
assert "Configured max size: 20 lines" in content
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_on_exceed_invalid_config(temp_dir, rootdir):
"""Test behavior with invalid configuration values."""
from sphinx.testing.util import SphinxTestApp
src_dir = rootdir / "basic"
app = SphinxTestApp(
srcdir=src_dir,
builddir=temp_dir,
buildername="html",
freshenv=True,
confoverrides={
"llms_txt_full_filename": "invalid-test.txt",
"llms_txt_full_max_size": 20,
"llms_txt_full_size_policy": "invalid_format", # Invalid config
},
)
app.build()
# Should fall back to default behavior (warn_skip)
output_file = Path(app.outdir) / "invalid-test.txt"
assert (
not output_file.exists()
), f"Output file {output_file} should not exist with invalid config fallback"
# Cleanup
sys.path[:] = app._saved_path
_clean_up_global_state()
if hasattr(app, "docutils_conf_path") and app.docutils_conf_path.exists():
app.docutils_conf_path.unlink()
def test_title_override(temp_dir, rootdir): def test_title_override(temp_dir, rootdir):
"""Test that the title override works correctly.""" """Test that the title override works correctly."""
from sphinx.testing.util import SphinxTestApp from sphinx.testing.util import SphinxTestApp
+37
View File
@@ -41,6 +41,42 @@ def test_setup_returns_valid_dict():
assert "parallel_write_safe" in result assert "parallel_write_safe" in result
def test_builder_inited_with_disallowed_builder():
"""Test that disallowed builders do not trigger extension setup."""
import sphinx_llms_txt
# Reset global state
sphinx_llms_txt._manager = sphinx_llms_txt.LLMSFullManager()
sphinx_llms_txt._root_first_paragraph = ""
# Mock a Sphinx app with a disallowed builder
class MockBuilder:
name = "text" # Not in allowed list
class MockApp:
def __init__(self):
self.config_values = {}
self.connections = {}
self.builder = MockBuilder()
def add_config_value(self, name, default, rebuild):
self.config_values[name] = (default, rebuild)
def connect(self, event, handler):
self.connections[event] = handler
app = MockApp()
setup(app)
# Trigger builder-inited
builder_inited_handler = app.connections["builder-inited"]
builder_inited_handler(app)
# With disallowed builder, other events should NOT be connected
assert "doctree-resolved" not in app.connections
assert "build-finished" not in app.connections
def test_document_collector_initialization(): def test_document_collector_initialization():
"""Test initialization of DocumentCollector.""" """Test initialization of DocumentCollector."""
collector = DocumentCollector() collector = DocumentCollector()
@@ -762,6 +798,7 @@ def test_summary_default_uses_first_paragraph():
llms_txt_full_file = True llms_txt_full_file = True
llms_txt_full_filename = "llms-full.txt" llms_txt_full_filename = "llms-full.txt"
llms_txt_full_max_size = None llms_txt_full_max_size = None
llms_txt_full_size_policy = "warn_skip"
llms_txt_directives = [] llms_txt_directives = []
llms_txt_exclude = [] llms_txt_exclude = []
llms_txt_code_files = [] llms_txt_code_files = []