Compare commits

..
18 Commits
Author SHA1 Message Date
Jared DillardandGitHub 46c2dec254 Use first paragraph as summary by default (#22) 2025-06-22 21:00:40 -07:00
Jared DillardandGitHub b56d93d265 Support source file suffix detection (#21) 2025-06-22 19:28:54 -07:00
Jared Dillard 5db5395889 add downloads badge to docs 2025-05-20 01:38:14 -07:00
Jared Dillard 04b0657dc5 add more badges 2025-05-20 01:35:40 -07:00
Jared Dillard 935a964c7e bump version to 0.2.3 2025-05-20 01:20:36 -07:00
Jared Dillard 51f6c71de3 update changelog 2025-05-20 01:18:44 -07:00
Jared DillardandGitHub 70defd3996 Remove get_and_resolve_toctree method (#19) 2025-05-20 01:08:40 -07:00
Jared DillardandGitHub 9ae05c6c13 Simplify _sources lookup (#18) 2025-05-20 00:50:06 -07:00
Jared Dillard 5581979cac make a href 2025-05-18 21:24:14 -07:00
Jared Dillard f70f1a26ec add soft transfer 2025-05-18 21:21:34 -07:00
Jared DillardandGitHub ed50138ae4 Update README.md 2025-05-18 21:12:16 -07:00
Jared DillardandGitHub 8f4d2c07c6 Update docs and README (#17)
* Move readme content to index.rst

* Clean up project name

* Add advanced configuration
2025-05-18 21:10:49 -07:00
Jared Dillard da7ee68076 support multi-line summaries 2025-05-18 18:17:30 -07:00
Jared Dillard 8ed31f13c4 strip whitespace from summary 2025-05-18 18:11:20 -07:00
Jared Dillard cc38abc8f2 add example links in docs 2025-05-18 17:57:09 -07:00
Jared Dillard bf368670db install pypi version 2025-05-18 17:51:58 -07:00
Jared Dillard a5cbdf15fa rename rtd config file 2025-05-18 17:36:15 -07:00
Jared DillardandGitHub e661ba3da7 Add sphinx docs (#16) 2025-05-18 17:31:00 -07:00
17 changed files with 1305 additions and 211 deletions
+15
View File
@@ -0,0 +1,15 @@
version: 2
build:
os: "ubuntu-20.04"
tools:
python: "3.10"
sphinx:
configuration: docs/source/conf.py
python:
install:
- requirements: docs/requirements.txt
- method: pip
path: .
+22
View File
@@ -1,6 +1,28 @@
Changelog Changelog
========= =========
0.3.0
-----
- Use first paragraph as default for ``llms_txt_summary``
`#22 <https://github.com/jdillard/sphinx-llms-txt/pull/22>`_
0.2.4
-----
- Support source file suffix detection
`#21 <https://github.com/jdillard/sphinx-llms-txt/pull/21>`_
0.2.3
-----
- Remove ``get_and_resolve_toctree`` method
`#19 <https://github.com/jdillard/sphinx-llms-txt/pull/19>`_
- Simplify ``_sources`` lookup
`#18 <https://github.com/jdillard/sphinx-llms-txt/pull/18>`_
- Add sphinx docs
`#16 <https://github.com/jdillard/sphinx-llms-txt/pull/16>`_
0.2.2 0.2.2
----- -----
+4 -81
View File
@@ -1,91 +1,14 @@
# Sphinx llms.txt generator # Sphinx llms.txt generator
A Sphinx extension that generates a summary `llms.txt` file, written in Markdown, and a single combined documentation `llms-full.txt` file, written in reStructuredText. A Sphinx extension that generates a summary `llms.txt` file and a single combined documentation `llms-full.txt` file.
[![PyPI version](https://img.shields.io/pypi/v/sphinx-llms-txt.svg)](https://pypi.python.org/pypi/sphinx-llms-txt) [![PyPI version](https://img.shields.io/pypi/v/sphinx-llms-txt.svg)](https://pypi.python.org/pypi/sphinx-llms-txt)
[![Downloads](https://static.pepy.tech/badge/sphinx-llms-txt/month)](https://pepy.tech/project/sphinx-llms-txt) [![Downloads](https://static.pepy.tech/badge/sphinx-llms-txt/month)](https://pepy.tech/project/sphinx-llms-txt)
[![Parallel Safe](https://img.shields.io/badge/parallel%20safe-true-brightgreen)](#)
## Installation ## Documentation
```bash See [sphinx-llms-txt documentation](https://sphinx-llms-txt.readthedocs.io/en/latest/index.html) for installation and configuration instructions.
pip install sphinx-llms-txt
```
## Usage
1. Add the extension to your Sphinx configuration (`conf.py`):
```python
extensions = [
'sphinx_llms_txt',
]
```
## Configuration Options
### `llms_txt_full_file`
- **Type**: boolean
- **Default**: `'True'`
- **Description**: Whether to write the single output file
### `llms_txt_full_filename`
- **Type**: string
- **Default**: `'llms-full.txt'`
- **Description**: Name of the single output file
### `llms_txt_full_max_size`
- **Type**: integer or `None`
- **Default**: `None` (no limit)
- **Description**: Sets a maximum line count for `llms_txt_full_filename`.
If exceeded, the file is skipped and a warning is shown, but the build still completes.
### `llms_txt_file`
- **Type**: boolean
- **Default**: `True`
- **Description**: Whether to write the summary information file
### `llms_txt_filename`
- **Type**: string
- **Default**: `llms.txt`
- **Description**: Name of the summary information file
### `llms_txt_directives`
- **Type**: list of strings
- **Default**: `[]`
- **Description**: List of custom directive names to process for path resolution.
### `llms_txt_title`
- **Type**: string or `None`
- **Default**: `None`
- **Description**: Overrides the Sphinx project name as the heading in `llms.txt`.
### `llms_txt_summary`
- **Type**: string or `None`
- **Default**: `None`
- **Description**: Optional, but recommended, summary description for `llms.txt`.
### `llms_txt_exclude`
- **Type**: list of strings
- **Default**: `[]`
- **Description**: A list of pages to ignore (e.g., `["page1", "page_with_*"]`).
## Features
- Creates `llms.txt` and `llms-full.txt`
- Automatically add content from `include` directives
- Resolves relative paths in directives like `image` and `figure` to use full paths
- Ability to add list of custom directives with `llms_txt_directives`
- Optionally, prepend a base URL using Sphinx's `html_baseurl`
- Ability to exclude pages
## License ## License
+20
View File
@@ -0,0 +1,20 @@
# Minimal makefile for Sphinx documentation
#
# You can set these variables from the command line.
SPHINXOPTS =
SPHINXBUILD = sphinx-build
SPHINXPROJ = SphinxLLMsTxt
SOURCEDIR = source
BUILDDIR = _build
# Put it first so that "make" without argument is like "make help".
help:
@$(SPHINXBUILD) -M help "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
.PHONY: help Makefile
# Catch-all target: route all unknown targets to Sphinx using the new
# "make mode" option. $(O) is meant as a shortcut for $(SPHINXOPTS).
%: Makefile
@$(SPHINXBUILD) -M $@ "$(SOURCEDIR)" "$(BUILDDIR)" $(SPHINXOPTS) $(O)
+6
View File
@@ -0,0 +1,6 @@
furo
esbonio
sphinx-contributors
sphinx
sphinx-llms-txt
sphinxext-opengraph
+170
View File
@@ -0,0 +1,170 @@
Advanced Configuration
======================
This page covers advanced configuration options for the sphinx-llms-txt extension.
.. _customizing_llms_files:
Customizing the LLMs Files
^^^^^^^^^^^^^^^^^^^^^^^^^^
By default, the extension generates two files:
1. ``llms.txt`` - A summary file in Markdown format
2. ``llms-full.txt`` - A complete documentation file in reStructuredText format
You can customize these files in several ways:
.. _changing_filenames:
Changing Filenames
~~~~~~~~~~~~~~~~~~
You can change the default filenames by setting these values in your ``conf.py``:
.. code-block:: python
llms_txt_filename = "custom-summary.txt"
llms_txt_full_filename = "custom-docs.txt"
.. _disabling_file_generation:
Disabling File Generation
~~~~~~~~~~~~~~~~~~~~~~~~~
If you only want one of the files, you can disable generation of the other:
.. code-block:: python
# Disable summary file
llms_txt_file = False
# Disable full documentation file
llms_txt_full_file = False
.. _custom_summary:
Adding a Custom Summary
~~~~~~~~~~~~~~~~~~~~~~~
The summary file can include a custom description of your project:
.. code-block:: python
llms_txt_summary = """
This documentation explains how to use MyProject to build amazing
applications. The project provides a comprehensive API for handling
data processing and visualization.
"""
.. note:: The summary can span multiple lines and will be properly formatted in the output file.
.. _custom_title:
Custom Title
~~~~~~~~~~~~
By default, the project name from Sphinx is used as the title in ``llms.txt``. You can override this:
.. code-block:: python
llms_txt_title = "My Custom Project Documentation"
.. _handling_large_documentation:
Handling Large Documentation
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
For very large documentation sets, generating the full documentation file might exceed reasonable size limits.
You can set a maximum line count:
.. code-block:: python
llms_txt_full_max_size = 10000 # Maximum 10,000 lines
If the generated file would exceed this limit, the extension will skip its generation and show a warning, allowing the build to complete.
.. tip:: Use :ref:`excluding_content` to remove less relevant pages.
.. _custom_directive_handling:
Custom Directive Handling
^^^^^^^^^^^^^^^^^^^^^^^^^
.. _path_resolution:
Path Resolution
~~~~~~~~~~~~~~~
The extension resolves paths in the common directives ``[ 'image', 'figure']`` by default.
You can add custom directives to this list:
.. code-block:: python
llms_txt_directives = [
"my-custom-image-directive",
"another-directive-with-paths",
]
This ensures that paths in your custom directives are properly resolved in the generated files.
.. _excluding_content:
Excluding Content
^^^^^^^^^^^^^^^^^
You can exclude specific pages from being included in the generated files:
.. code-block:: python
llms_txt_exclude = [
"search", # Exclude the search page
"genindex", # Exclude the index page
"private_*", # Exclude all pages starting with 'private_'
]
This is useful for excluding auto-generated pages, indexes, or content that isn't relevant for LLM consumption.
.. _using_html_baseurl:
Using HTML Base URL
^^^^^^^^^^^^^^^^^^^
If you want to include absolute URLs for resources in your documentation, you can use Sphinx's built-in ``html_baseurl`` configuration:
.. code-block:: python
html_baseurl = "https://example.com/docs/"
When this option is set, all resolved paths in directives will be prefixed with this URL, creating absolute paths in the generated files.
.. _integration_examples:
Integration Examples
^^^^^^^^^^^^^^^^^^^^
Complete Configuration Example
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Here's a complete example showing multiple :doc:`configuration-values`:
.. code-block:: python
# File names and generation options
llms_txt_filename = "ai-summary.txt"
llms_txt_full_filename = "ai-full-docs.txt"
llms_txt_full_max_size = 50000
# Content customization
llms_txt_title = "Project Documentation for AI Assistants"
llms_txt_summary = """
This is a comprehensive documentation set for our project.
It includes API references, usage examples, and tutorials.
"""
# Path handling
html_baseurl = "https://docs.example.com/"
llms_txt_directives = ["custom-image", "custom-include"]
# Content filtering
llms_txt_exclude = ["search", "genindex", "404", "private_*"]
+1
View File
@@ -0,0 +1 @@
.. include:: ../../CHANGELOG.rst
+101
View File
@@ -0,0 +1,101 @@
#
# Configuration file for the Sphinx documentation builder.
#
# This file does only contain a selection of the most common options. For a
# full list see the documentation:
# http://www.sphinx-doc.org/en/master/config
# -- Path setup --------------------------------------------------------------
import re
import subprocess
# -- Project information -----------------------------------------------------
project = "sphinx-llms-txt"
copyright = "Jared Dillard"
author = "Jared Dillard"
llms_txt_summary = """
A Sphinx extension that generates a summary llms.txt file,written in Markdown,
and a single combined documentation llms-full.txt file, written in reStructuredText.
"""
# check if the current commit is tagged as a release (vX.Y.Z)
try:
GIT_TAG_OUTPUT = subprocess.check_output(["git", "tag", "--points-at", "HEAD"])
current_tag = GIT_TAG_OUTPUT.decode().strip()
if re.match(r"^v(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)$", current_tag):
version = current_tag
else:
version = "latest"
except (subprocess.CalledProcessError, FileNotFoundError):
version = "latest"
# The full version, including alpha/beta/rc tags
release = ""
# -- General configuration ---------------------------------------------------
# If your documentation needs a minimal Sphinx version, state it here.
#
# needs_sphinx = '1.0'
# Add any Sphinx extension module names here, as strings. They can be
# extensions coming with Sphinx (named 'sphinx.ext.*') or your custom
# ones.
extensions = [
"sphinx.ext.intersphinx",
"sphinx_contributors",
"sphinx_llms_txt",
]
# The language for content autogenerated by Sphinx. Refer to documentation
# for a list of supported languages.
#
# This is also used if you do content translation via gettext catalogs.
# Usually you set "language" from the command line for these cases.
language = "en"
# List of patterns, relative to source directory, that match files and
# directories to ignore when looking for source files.
# This pattern also affects html_static_path and html_extra_path.
exclude_patterns = ["_build", "Thumbs.db", ".DS_Store"]
# The name of the Pygments (syntax highlighting) style to use.
pygments_style = "sphinx"
intersphinx_mapping = {
"sphinx": ("https://www.sphinx-doc.org/en/master/", None),
}
# -- Options for HTML output -------------------------------------------------
# The theme to use for HTML and HTML Help pages. See the documentation for
# a list of builtin themes.
#
html_theme = "furo"
# Theme options are theme-specific and customize the look and feel of a theme
# further. For a list of options available for each theme, see the
# documentation.
#
html_theme_options = {}
html_baseurl = "https://sphinx-llms-txt.readthedocs.org/"
# -- Options for HTMLHelp output ---------------------------------------------
# Output file base name for HTML help builder.
htmlhelp_basename = "SphinxLLMsTxtdoc"
def setup(app):
app.add_object_type(
"confval",
"confval",
objname="configuration value",
indextemplate="pair: %s; configuration value",
)
+84
View File
@@ -0,0 +1,84 @@
Project Configuration Values
============================
.. confval:: llms_txt_full_file
- **Type**: boolean
- **Default**: ``True``
- **Description**: Whether to write the single output file.
See :ref:`disabling_file_generation`.
.. versionadded:: 0.1.0
.. confval:: llms_txt_full_filename
- **Type**: string
- **Default**: ``'llms-full.txt'``
- **Description**: Name of the single output file.
See :ref:`changing_filenames`.
.. versionadded:: 0.1.0
.. confval:: llms_txt_full_max_size
- **Type**: integer or ``None``
- **Default**: ``None`` (no limit)
- **Description**: Sets a maximum line count for ``llms_txt_full_filename``.
If exceeded, the file is skipped and a warning is shown, but the build still completes.
See :ref:`handling_large_documentation`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_file
- **Type**: boolean
- **Default**: ``True``
- **Description**: Whether to write the summary information file.
See :ref:`disabling_file_generation`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_filename
- **Type**: string
- **Default**: ``llms.txt``
- **Description**: Name of the summary information file.
See :ref:`changing_filenames`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_directives
- **Type**: list of strings
- **Default**: ``[]`` (empty list)
- **Description**: List of custom directive names to process for path resolution.
See :ref:`path_resolution`.
.. versionadded:: 0.1.0
.. confval:: llms_txt_title
- **Type**: string or ``None``
- **Default**: ``None``
- **Description**: Overrides the Sphinx project name as the heading in ``llms.txt``.
See :ref:`custom_title`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_summary
- **Type**: string
- **Default**: The first paragraph in the root document, else an empty string
- **Description**: Optional, but recommended, summary description for ``llms.txt``.
See :ref:`custom_summary`.
.. versionadded:: 0.2.0
.. confval:: llms_txt_exclude
- **Type**: list of strings
- **Default**: ``[]``
- **Description**: A list of pages to ignore.
See :ref:`excluding_content`.
.. versionadded:: 0.2.1
+48
View File
@@ -0,0 +1,48 @@
Contributing
============
You will need to set up a development environment to make and test your changes before submitting them.
Local development
-----------------
#. Clone the `sphinx-llms-txt repository`_.
#. Create and activate a virtual environment:
.. code-block:: console
python3 -m venv .venv
source .venv/bin/activate
#. Install development dependencies:
.. code-block:: console
pip install -e ".[dev]"
#. Install pre-commit Git hook scripts:
.. code-block:: console
pre-commit install
Testing changes
---------------
Run ``pytest`` before committing changes.
Current contributors
--------------------
Thanks to all who have contributed!
The people that have improved the code:
.. contributors:: jdillard/sphinx-llms-txt
:avatars:
:limit: 100
:exclude: pre-commit-ci[bot],dependabot[bot]
:order: ASC
.. _sphinx-llms-txt repository: https://github.com/jdillard/sphinx-llms-txt
+50
View File
@@ -0,0 +1,50 @@
Getting Started
===============
Demo
----
You can see this Sphinx project's `llms.txt`_ and `llms-full.txt`_ files as a simple example.
Installation
------------
Directly install via ``pip`` by using:
.. code-block:: bash
pip install sphinx-llms-txt
Usage
-----
Add the extension to your Sphinx configuration (``conf.py``):
.. code-block:: python
extensions = [
'sphinx_llms_txt',
]
Once added, the extension will automatically generate the LLMs.txt files during the build process.
See :doc:`advanced-configuration` for more information about how to use **sphinx-llms-txt**.
How It Works
------------
During the Sphinx build process:
1. **Content Collection**: Scans all of your documentation's ``_source`` pages and collects their content
2. **Directive Processing**: Resolves ``include`` directives by automatically incorporating their content
3. **Path Resolution**: Transforms relative paths in directives to full paths
4. **Output Generation**: Creates two optional files:
- ``llms.txt``: A concise summary of your documentation, in Markdown
- ``llms-full.txt``: A comprehensive version with all documentation content, in reStructuredText
5. **Content Filtering**: Allows you to exclude specific pages from the generated files
.. _llms.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms.txt
.. _llms-full.txt: https://sphinx-llms-txt.readthedocs.io/en/latest/llms-full.txt
+31
View File
@@ -0,0 +1,31 @@
Sphinx llms.txt Generator
=========================
A `Sphinx`_ extension that generates a summary ``llms.txt`` file, written in Markdown, and a single combined documentation ``llms-full.txt`` file, written in reStructuredText.
|PyPI version| |Downloads| |Parallel Safe| |GitHub Stars|
.. toctree::
:maxdepth: 2
getting-started
advanced-configuration
configuration-values
contributing
changelog
.. _Sphinx: http://sphinx-doc.org/
.. |PyPI version| image:: https://img.shields.io/pypi/v/sphinx-llms-txt.svg
:target: https://pypi.python.org/pypi/sphinx-llms-txt
:alt: Latest PyPi Version
.. |Downloads| image:: https://static.pepy.tech/badge/sphinx-llms-txt/month
:target: https://pepy.tech/project/sphinx-llms-txt
:alt: PyPi Downloads per month
.. |Parallel Safe| image:: https://img.shields.io/badge/parallel%20safe-true-brightgreen
:target: #
:alt: Parallel read/write safe
.. |GitHub Stars| image:: https://img.shields.io/github/stars/jdillard/sphinx-llms-txt?style=social
:target: https://github.com/jdillard/sphinx-llms-txt
:alt: GitHub Repository stars
+23 -4
View File
@@ -12,7 +12,7 @@ from .manager import LLMSFullManager
from .processor import DocumentProcessor from .processor import DocumentProcessor
from .writer import FileWriter from .writer import FileWriter
__version__ = "0.2.2" __version__ = "0.3.0"
# Export classes needed by tests # Export classes needed by tests
__all__ = [ __all__ = [
@@ -25,9 +25,14 @@ __all__ = [
# Global manager instance # Global manager instance
_manager = LLMSFullManager() _manager = LLMSFullManager()
# Store root document first paragraph
_root_first_paragraph = ""
def doctree_resolved(app: Sphinx, doctree, docname: str): def doctree_resolved(app: Sphinx, doctree, docname: str):
"""Called when a docname has been resolved to a document.""" """Called when a docname has been resolved to a document."""
global _root_first_paragraph
# Extract title from the document # Extract title from the document
title = None title = None
# findall() returns a generator, convert to list to check if it has elements # findall() returns a generator, convert to list to check if it has elements
@@ -38,6 +43,14 @@ def doctree_resolved(app: Sphinx, doctree, docname: str):
if title: if title:
_manager.update_page_title(docname, title) _manager.update_page_title(docname, title)
# Extract first paragraph from root document
if docname == app.config.master_doc:
for node in doctree.traverse(nodes.paragraph):
first_para = node.astext()
if first_para:
_root_first_paragraph = first_para
break
def build_finished(app: Sphinx, exception): def build_finished(app: Sphinx, exception):
"""Called when the build is finished.""" """Called when the build is finished."""
@@ -47,12 +60,17 @@ def build_finished(app: Sphinx, exception):
_manager.set_master_doc(app.config.master_doc) _manager.set_master_doc(app.config.master_doc)
_manager.set_app(app) _manager.set_app(app)
# Get the summary - use configured value or extracted first paragraph
summary = app.config.llms_txt_summary
if summary is None:
summary = _root_first_paragraph
# Set up configuration # Set up configuration
config = { config = {
"llms_txt_file": app.config.llms_txt_file, "llms_txt_file": app.config.llms_txt_file,
"llms_txt_filename": app.config.llms_txt_filename, "llms_txt_filename": app.config.llms_txt_filename,
"llms_txt_title": app.config.llms_txt_title, "llms_txt_title": app.config.llms_txt_title,
"llms_txt_summary": app.config.llms_txt_summary, "llms_txt_summary": summary,
"llms_txt_full_file": app.config.llms_txt_full_file, "llms_txt_full_file": app.config.llms_txt_full_file,
"llms_txt_full_filename": app.config.llms_txt_full_filename, "llms_txt_full_filename": app.config.llms_txt_full_filename,
"llms_txt_full_max_size": app.config.llms_txt_full_max_size, "llms_txt_full_max_size": app.config.llms_txt_full_max_size,
@@ -91,9 +109,10 @@ def setup(app: Sphinx) -> Dict[str, Any]:
app.connect("doctree-resolved", doctree_resolved) app.connect("doctree-resolved", doctree_resolved)
app.connect("build-finished", build_finished) app.connect("build-finished", build_finished)
# Reset manager for each build # Reset manager and root paragraph for each build
global _manager global _manager, _root_first_paragraph
_manager = LLMSFullManager() _manager = LLMSFullManager()
_root_first_paragraph = ""
return { return {
"version": __version__, "version": __version__,
+119 -26
View File
@@ -3,7 +3,7 @@ Document collector module for sphinx-llms-txt.
""" """
import fnmatch import fnmatch
from typing import Any, Dict, List from typing import Any, Dict, List, Tuple
from sphinx.environment import BuildEnvironment from sphinx.environment import BuildEnvironment
from sphinx.util import logging from sphinx.util import logging
@@ -19,6 +19,7 @@ class DocumentCollector:
self.master_doc: str = None self.master_doc: str = None
self.env: BuildEnvironment = None self.env: BuildEnvironment = None
self.config: Dict[str, Any] = {} self.config: Dict[str, Any] = {}
self.app = None
def set_master_doc(self, master_doc: str): def set_master_doc(self, master_doc: str):
"""Set the master document name.""" """Set the master document name."""
@@ -37,8 +38,73 @@ class DocumentCollector:
"""Set configuration options.""" """Set configuration options."""
self.config = config self.config = config
def get_page_order(self) -> List[str]: def set_app(self, app):
"""Get the correct page order from the toctree structure.""" """Set the Sphinx application reference."""
self.app = app
def _get_source_suffixes(self):
"""Get all valid source file suffixes from Sphinx configuration.
Returns:
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
"""
if not self.app:
return [".rst"] # Default fallback
source_suffix = self.app.config.source_suffix
if isinstance(source_suffix, dict):
return list(source_suffix.keys())
elif isinstance(source_suffix, list):
return source_suffix
else:
return [source_suffix] # String format
def _get_docname_suffix(self, docname: str, sources_dir) -> str:
"""
Determine the source suffix for a given docname by checking which
file exists.
Args:
docname: The document name to check
sources_dir: Path to the _sources directory
Returns:
The source suffix if found, or None if no matching file exists
"""
if not sources_dir or not sources_dir.exists():
return None
# Get the source link suffix from Sphinx config
source_link_suffix = ""
if self.app and hasattr(self.app.config, "html_sourcelink_suffix"):
source_link_suffix = self.app.config.html_sourcelink_suffix
# Handle empty string case specially
if source_link_suffix == "":
source_link_suffix = "" # Keep it empty
elif not source_link_suffix.startswith("."):
source_link_suffix = "." + source_link_suffix
# Get the source file suffixes from Sphinx config
source_suffixes = self._get_source_suffixes()
# Try to find the source file with any of the valid source suffixes
for src_suffix in source_suffixes:
candidate_file = sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
if candidate_file.exists():
return src_suffix
return None
def get_page_order(self, sources_dir=None) -> List[Tuple[str, str]]:
"""Get the correct page order from the toctree structure.
Args:
sources_dir: Optional path to _sources directory for suffix detection
Returns:
List of tuples (docname, source_suffix) in toctree order
"""
if not self.env or not self.master_doc: if not self.env or not self.master_doc:
return [] return []
@@ -52,9 +118,12 @@ class DocumentCollector:
visited.add(docname) visited.add(docname)
# Add the current document # Add the current document with its suffix
if docname not in page_order: if docname not in [doc for doc, _ in page_order]:
page_order.append(docname) suffix = None
if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix))
# Check for toctree entries in this document # Check for toctree entries in this document
try: try:
@@ -65,21 +134,34 @@ class DocumentCollector:
): ):
for child_docname in self.env.toctree_includes[docname]: for child_docname in self.env.toctree_includes[docname]:
collect_from_toctree(child_docname) collect_from_toctree(child_docname)
else: # Try to use dependencies to find related documents
# Fallback: try to resolve and parse the toctree elif (
toctree = self.env.get_and_resolve_toctree(docname, None) hasattr(self.env, "dependencies")
if toctree: and docname in self.env.dependencies
from docutils import nodes ):
# Extract the dependent documents from the dependencies dict
for child_docname in self.env.dependencies[docname]:
# Only add documents actually in the document set
if (
hasattr(self.env, "all_docs")
and child_docname in self.env.all_docs
):
collect_from_toctree(child_docname)
# Fallback to titles or other available references
elif hasattr(self.env, "titles") and hasattr(self.env, "all_docs"):
# Get all document names
all_docnames = list(self.env.all_docs.keys())
for node in list(toctree.findall(nodes.reference)): # Look for documents that might be related (have similar paths)
if "refuri" in node.attributes: current_prefix = "/".join(docname.split("/")[:-1])
refuri = node.attributes["refuri"] if current_prefix:
if refuri and refuri.endswith(".html"): for child_docname in all_docnames:
child_docname = refuri[:-5] # Remove .html # Documents in the same directory might be related
if ( if (
child_docname != docname child_docname.startswith(current_prefix)
): # Avoid circular references and child_docname != docname
collect_from_toctree(child_docname) ):
collect_from_toctree(child_docname)
except Exception as e: except Exception as e:
logger.debug(f"Could not get toctree for {docname}: {e}") logger.debug(f"Could not get toctree for {docname}: {e}")
@@ -88,22 +170,33 @@ class DocumentCollector:
# Add any remaining documents not in the toctree (sorted) # Add any remaining documents not in the toctree (sorted)
if hasattr(self.env, "all_docs"): if hasattr(self.env, "all_docs"):
processed_docnames = {doc for doc, _ in page_order}
remaining = sorted( remaining = sorted(
[doc for doc in self.env.all_docs.keys() if doc not in page_order] [
doc
for doc in self.env.all_docs.keys()
if doc not in processed_docnames
]
) )
page_order.extend(remaining) for docname in remaining:
suffix = None
if sources_dir:
suffix = self._get_docname_suffix(docname, sources_dir)
page_order.append((docname, suffix))
return page_order return page_order
def filter_excluded_pages(self, page_order: List[str]) -> List[str]: def filter_excluded_pages(
self, page_order: List[Tuple[str, str]]
) -> List[Tuple[str, str]]:
"""Filter out excluded pages from the page order.""" """Filter out excluded pages from the page order."""
exclude_patterns = self.config.get("llms_txt_exclude") exclude_patterns = self.config.get("llms_txt_exclude")
if exclude_patterns: if exclude_patterns:
return [ return [
page (docname, suffix)
for page in page_order for docname, suffix in page_order
if not any( if not any(
self._match_exclude_pattern(page, pattern) self._match_exclude_pattern(docname, pattern)
for pattern in exclude_patterns for pattern in exclude_patterns
) )
] ]
+119 -94
View File
@@ -56,6 +56,7 @@ class LLMSFullManager:
def set_app(self, app: Sphinx): def set_app(self, app: Sphinx):
"""Set the Sphinx application reference.""" """Set the Sphinx application reference."""
self.app = app self.app = app
self.collector.set_app(app)
if self.writer: if self.writer:
self.writer.app = app self.writer.app = app
@@ -69,23 +70,7 @@ class LLMSFullManager:
self.processor = DocumentProcessor(self.config, srcdir) self.processor = DocumentProcessor(self.config, srcdir)
self.writer = FileWriter(self.config, outdir, self.app) self.writer = FileWriter(self.config, outdir, self.app)
# Get the correct page order # Find sources directory first so we can pass it to get_page_order
page_order = self.collector.get_page_order()
if not page_order:
logger.warning(
"Could not determine page order, skipping llms-full creation"
)
return
# Apply exclusion filter if configured
page_order = self.collector.filter_excluded_pages(page_order)
# Determine output file name and location
output_filename = self.config.get("llms_txt_full_filename")
output_path = Path(outdir) / output_filename
# Find sources directory
sources_dir = None sources_dir = None
possible_sources = [ possible_sources = [
Path(outdir) / "_sources", Path(outdir) / "_sources",
@@ -104,14 +89,23 @@ class LLMSFullManager:
) )
return return
# Collect all available source files # Get the correct page order with source suffixes
txt_files = {} page_order = self.collector.get_page_order(sources_dir)
for f in sources_dir.glob("**/*.txt"):
logger.debug(f"sphinx-llms-txt: Found source file: {f.stem} at {f}") if not page_order:
txt_files[f.stem] = f logger.warning(
"Could not determine page order, skipping llms-full creation"
)
return
# Apply exclusion filter if configured
page_order = self.collector.filter_excluded_pages(page_order)
# Determine output file name and location
output_filename = self.config.get("llms_txt_full_filename")
output_path = Path(outdir) / output_filename
# Log discovered files and page order # Log discovered files and page order
logger.debug(f"sphinx-llms-txt: Found {len(txt_files)} source files")
logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}") logger.debug(f"sphinx-llms-txt: Page order (after exclusion): {page_order}")
# Log exclusion patterns # Log exclusion patterns
@@ -119,33 +113,43 @@ class LLMSFullManager:
if exclude_patterns: if exclude_patterns:
logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}") logger.debug(f"sphinx-llms-txt: Exclusion patterns: {exclude_patterns}")
# Create a mapping from docnames to actual file names # Create a mapping from docnames to source files
docname_to_file = {} docname_to_file = {}
# Try exact matches first # Get the source link suffix from Sphinx config
for docname in page_order: source_link_suffix = (
self.app.config.html_sourcelink_suffix if self.app else ".txt"
)
# Handle empty string case specially
if source_link_suffix == "":
source_link_suffix = "" # Keep it empty
elif not source_link_suffix.startswith("."):
source_link_suffix = "." + source_link_suffix
# Process each (docname, suffix) in the page order
for docname, src_suffix in page_order:
# Skip excluded pages # Skip excluded pages
if any( if exclude_patterns and any(
self.collector._match_exclude_pattern(docname, pattern) self.collector._match_exclude_pattern(docname, pattern)
for pattern in exclude_patterns for pattern in exclude_patterns
): ):
continue continue
if docname in txt_files: # Build the source file path directly using the known suffix
docname_to_file[docname] = txt_files[docname] if src_suffix:
source_file = sources_dir / f"{docname}{src_suffix}{source_link_suffix}"
if source_file.exists():
docname_to_file[docname] = source_file
else:
logger.warning(
f"sphinx-llms-txt: Source file not found for: {docname}."
f"Expected: {docname}{src_suffix}{source_link_suffix}"
)
else: else:
# Try with .rst extension logger.warning(
if f"{docname}.rst" in txt_files: f"sphinx-llms-txt: No source suffix determined for: {docname}"
docname_to_file[docname] = txt_files[f"{docname}.rst"] )
# Try with .txt extension
elif f"{docname}.txt" in txt_files:
docname_to_file[docname] = txt_files[f"{docname}.txt"]
# Try with underscores instead of hyphens
elif docname.replace("-", "_") in txt_files:
docname_to_file[docname] = txt_files[docname.replace("-", "_")]
# Try with hyphens instead of underscores
elif docname.replace("_", "-") in txt_files:
docname_to_file[docname] = txt_files[docname.replace("_", "-")]
# Generate content # Generate content
content_parts = [] content_parts = []
@@ -156,7 +160,7 @@ class LLMSFullManager:
max_lines = self.config.get("llms_txt_full_max_size") max_lines = self.config.get("llms_txt_full_max_size")
abort_due_to_max_lines = False abort_due_to_max_lines = False
for docname in page_order: for docname, _ in page_order:
if docname in docname_to_file: if docname in docname_to_file:
file_path = docname_to_file[docname] file_path = docname_to_file[docname]
content, line_count = self._read_source_file(file_path, docname) content, line_count = self._read_source_file(file_path, docname)
@@ -190,65 +194,68 @@ class LLMSFullManager:
added_files.add(file_path.stem) added_files.add(file_path.stem)
total_line_count += line_count total_line_count += line_count
else: else:
logger.warning(f"sphinx-llm-txt: Source file not found for: {docname}") logger.warning(
f"sphinx-llms-txt: Source file not found for: {docname}. Check that"
f" file exists at _sources/{docname}[suffix]{source_link_suffix}"
)
# Add any remaining files (in alphabetical order) if not aborted # Add any remaining files (in alphabetical order) that aren't in the page order
if not abort_due_to_max_lines: if not abort_due_to_max_lines:
# Apply the same exclusion filter to remaining files # Get all source files in the _sources directory using configured suffixes
exclude_patterns = self.config.get("llms_txt_exclude") source_suffixes = self._get_source_suffixes()
all_source_files = []
for src_suffix in source_suffixes:
glob_pattern = f"**/*{src_suffix}{source_link_suffix}"
all_source_files.extend(sources_dir.glob(glob_pattern))
# Create a set of files to exclude based on their basename processed_paths = set(file.resolve() for file in docname_to_file.values())
excluded_files = set()
for pattern in exclude_patterns:
if "*" not in pattern and "?" not in pattern:
# For exact patterns, add variants
excluded_files.add(pattern)
excluded_files.add(f"{pattern}.rst")
excluded_files.add(f"{pattern}.txt")
excluded_files.add(pattern.replace("-", "_"))
excluded_files.add(pattern.replace("_", "-"))
# Filter remaining files # Find files that haven't been processed yet
remaining_files = sorted( remaining_source_files = [
[ f for f in all_source_files if f.resolve() not in processed_paths
name ]
for name in txt_files
if name not in added_files # Sort the remaining files for consistent ordering
and name not in excluded_files remaining_source_files.sort()
and not any(
self.collector._match_exclude_pattern(name, pattern) if remaining_source_files:
for pattern in exclude_patterns logger.info(
) f"Found {len(remaining_source_files)} additional files not in"
] f" toctree"
) )
if remaining_files:
logger.info(f"Adding remaining files: {remaining_files}") for file_path in remaining_source_files:
for file_stem in remaining_files: # Extract docname from path by removing the source and link suffixes
file_path = txt_files[file_stem] rel_path = str(file_path.relative_to(sources_dir))
content, line_count = self._read_source_file(file_path, file_stem) docname = None
# Try each source suffix to find which one this file uses
for src_suffix in source_suffixes:
combined_suffix = f"{src_suffix}{source_link_suffix}"
if rel_path.endswith(combined_suffix):
docname = rel_path[: -len(combined_suffix)] # Remove suffix
break
if docname is None:
continue
# Skip excluded docnames
if exclude_patterns and any(
self.collector._match_exclude_pattern(docname, pattern)
for pattern in exclude_patterns
):
logger.debug(f"sphinx-llms-txt: Skipping excluded file: {docname}")
continue
# Read and process the file
content, line_count = self._read_source_file(file_path, docname)
# Check if adding this file would exceed the maximum line count # Check if adding this file would exceed the maximum line count
if max_lines is not None and total_line_count + line_count > max_lines: if max_lines is not None and total_line_count + line_count > max_lines:
break break
# Double-check that this file should be included if content:
should_include = True logger.debug(f"sphinx-llms-txt: Adding remaining file: {docname}")
file_stem = file_path.stem
exclude_patterns = self.config.get("llms_txt_exclude")
if exclude_patterns:
# Check stem against exclusion patterns
if any(
self.collector._match_exclude_pattern(file_stem, pattern)
for pattern in exclude_patterns
):
logger.debug(
"sphinx-llms-txt: Final exclusion check removed remaining"
f" file: {file_stem}"
)
should_include = False
if content and should_include:
content_parts.append(content) content_parts.append(content)
total_line_count += line_count total_line_count += line_count
@@ -258,7 +265,7 @@ class LLMSFullManager:
max_lines is not None and total_line_count > max_lines max_lines is not None and total_line_count > max_lines
): ):
logger.warning( logger.warning(
f"sphinx-llm-txt: Max line limit ({max_lines}) exceeded:" f"sphinx-llms-txt: Max line limit ({max_lines}) exceeded:"
f" {total_line_count} > {max_lines}. " f" {total_line_count} > {max_lines}. "
f"Not creating llms-full.txt file." f"Not creating llms-full.txt file."
) )
@@ -325,5 +332,23 @@ class LLMSFullManager:
return content_str, line_count + 1 return content_str, line_count + 1
except Exception as e: except Exception as e:
logger.error(f"sphinx-llm-txt: Error reading source file {file_path}: {e}") logger.error(f"sphinx-llms-txt: Error reading source file {file_path}: {e}")
return "", 0 return "", 0
def _get_source_suffixes(self):
"""Get all valid source file suffixes from Sphinx configuration.
Returns:
list: List of source file suffixes (e.g., ['.rst', '.md', '.txt'])
"""
if not self.app:
return [".rst"] # Default fallback
source_suffix = self.app.config.source_suffix
if isinstance(source_suffix, dict):
return list(source_suffix.keys())
elif isinstance(source_suffix, list):
return source_suffix
else:
return [source_suffix] # String format
+18 -6
View File
@@ -3,7 +3,7 @@ File writer module for sphinx-llms-txt.
""" """
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List from typing import Any, Dict, List, Tuple, Union
from sphinx.application import Sphinx from sphinx.application import Sphinx
from sphinx.util import logging from sphinx.util import logging
@@ -42,19 +42,19 @@ class FileWriter:
) )
return True return True
except Exception as e: except Exception as e:
logger.error(f"sphinx-llm-txt: Error writing combined sources file: {e}") logger.error(f"sphinx-llms-txt: Error writing combined sources file: {e}")
return False return False
def write_verbose_info_to_file( def write_verbose_info_to_file(
self, self,
page_order: List[str], page_order: Union[List[str], List[Tuple[str, str]]],
page_titles: Dict[str, str], page_titles: Dict[str, str],
total_line_count: int = 0, total_line_count: int = 0,
) -> bool: ) -> bool:
"""Write summary information to the llms.txt file. """Write summary information to the llms.txt file.
Args: Args:
page_order: Ordered list of document names page_order: Ordered list of document names or (docname, suffix) tuples
page_titles: Dictionary mapping docnames to titles page_titles: Dictionary mapping docnames to titles
total_line_count: Total number of lines in the combined content total_line_count: Total number of lines in the combined content
@@ -86,7 +86,14 @@ class FileWriter:
# Add description if available # Add description if available
description = self.config.get("llms_txt_summary", "") description = self.config.get("llms_txt_summary", "")
if description: if description:
f.write(f"> {description}\n\n") # Trim leading and trailing whitespace
description = description.strip()
if description:
# Only add blockquote if description is not empty
# Replace newlines with newline + blockquote marker to maintain
# blockquote formatting
description = description.replace("\n", "\n> ")
f.write(f"> {description}\n\n")
f.write("## Docs\n\n") f.write("## Docs\n\n")
# Get base URL from config # Get base URL from config
@@ -95,7 +102,12 @@ class FileWriter:
if not base_url.endswith("/"): if not base_url.endswith("/"):
base_url += "/" base_url += "/"
for docname in page_order: for item in page_order:
# Handle both old format (str) and new format (tuple)
if isinstance(item, tuple):
docname, _ = item
else:
docname = item
title = page_titles.get(docname, docname) title = page_titles.get(docname, docname)
f.write(f"- [{title}]({base_url}{docname}.html)\n") f.write(f"- [{title}]({base_url}{docname}.html)\n")
+474
View File
@@ -334,3 +334,477 @@ def test_write_verbose_info_with_baseurl(tmp_path):
assert "- [Home Page](https://example.org/index.html)" in content assert "- [Home Page](https://example.org/index.html)" in content
assert "- [About Us](https://example.org/about.html)" in content assert "- [About Us](https://example.org/about.html)" in content
def test_get_source_suffixes_with_dict():
"""Test _get_source_suffixes method with dict source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with dict source_suffix
class MockApp:
class Config:
source_suffix = {".rst": None, ".md": None, ".txt": None}
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert set(suffixes) == {".rst", ".md", ".txt"}
def test_get_source_suffixes_with_list():
"""Test _get_source_suffixes method with list source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with list source_suffix
class MockApp:
class Config:
source_suffix = [".rst", ".md"]
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst", ".md"]
def test_get_source_suffixes_with_string():
"""Test _get_source_suffixes method with string source_suffix."""
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with string source_suffix
class MockApp:
class Config:
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst"]
def test_get_source_suffixes_no_app():
"""Test _get_source_suffixes method with no app set."""
from sphinx_llms_txt.manager import LLMSFullManager
manager = LLMSFullManager()
suffixes = manager._get_source_suffixes()
assert suffixes == [".rst"] # Default fallback
def test_html_sourcelink_suffix_default():
"""Test html_sourcelink_suffix defaults to .txt when no app is set."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
manager = LLMSFullManager()
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with default .txt suffix
test_file = f"{sources_dir}/index.rst.txt"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it uses .txt as the default suffix
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed (check if output file exists)
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_html_sourcelink_suffix_custom():
"""Test html_sourcelink_suffix uses custom value from Sphinx config."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with custom html_sourcelink_suffix
class MockApp:
class Config:
html_sourcelink_suffix = "source"
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with custom .source suffix
test_file = f"{sources_dir}/index.rst.source"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it uses .source as the custom suffix
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_html_sourcelink_suffix_with_dot():
"""Test html_sourcelink_suffix adds dot if missing."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with html_sourcelink_suffix without leading dot
class MockApp:
class Config:
html_sourcelink_suffix = "src" # No leading dot
source_suffix = ".rst"
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create a test source file with .src suffix (dot should be added automatically)
test_file = f"{sources_dir}/index.rst.src"
with open(test_file, "w") as f:
f.write("Test content")
# Mock env with minimal required attributes
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Test Title"})()
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test that it adds the dot and finds the file
manager.combine_sources(outdir, srcdir)
# Verify the file was found and processed
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
def test_mixed_source_file_formats():
"""Test handling of mixed source file formats (.rst, .md, .txt)."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with multiple source suffixes
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
source_suffix = {".rst": None, ".md": None, ".txt": None}
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create test source files with different formats
files_to_create = [
f"{sources_dir}/page1.rst.txt",
f"{sources_dir}/page2.md.txt",
f"{sources_dir}/page3.txt.txt",
]
for test_file in files_to_create:
with open(test_file, "w") as f:
f.write(f"Content for {os.path.basename(test_file)}")
# Mock env with all documents
class MockEnv:
all_docs = {"page1": None, "page2": None, "page3": None}
titles = {
"page1": type("TitleNode", (), {"astext": lambda: "Page 1"})(),
"page2": type("TitleNode", (), {"astext": lambda: "Page 2"})(),
"page3": type("TitleNode", (), {"astext": lambda: "Page 3"})(),
}
toctree_includes = {}
manager.set_env(MockEnv())
manager.set_master_doc("page1")
# Test that all file formats are found and processed
manager.combine_sources(outdir, srcdir)
# Verify the output file was created and contains content from all formats
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
with open(output_file, "r") as f:
content = f.read()
# Should contain content from all three files
assert "Content for page1.rst.txt" in content
assert "Content for page2.md.txt" in content
assert "Content for page3.txt.txt" in content
def test_source_suffix_detection_priority():
"""Test source suffix detection tries formats in correct order for docnames."""
import tempfile
from sphinx_llms_txt.manager import LLMSFullManager
# Mock Sphinx app with ordered source suffixes
class MockApp:
class Config:
html_sourcelink_suffix = ".txt"
source_suffix = [".rst", ".md"] # rst has priority over md
config = Config()
manager = LLMSFullManager()
manager.set_app(MockApp())
manager.set_config(
{
"llms_txt_full_filename": "test.txt",
"llms_txt_exclude": [],
"llms_txt_directives": [],
}
)
# Create a temporary directory structure
with tempfile.TemporaryDirectory() as tmpdir:
outdir = f"{tmpdir}/build"
srcdir = f"{tmpdir}/source"
sources_dir = f"{outdir}/_sources"
# Create directories
import os
os.makedirs(sources_dir, exist_ok=True)
os.makedirs(srcdir, exist_ok=True)
# Create both .rst and .md versions of the same document
# Only create files for the specific docname "index"
rst_file = f"{sources_dir}/index.rst.txt"
md_file = f"{sources_dir}/index.md.txt"
with open(rst_file, "w") as f:
f.write("RST content for index")
with open(md_file, "w") as f:
f.write("Markdown content for index")
# Mock env with only the index document
class MockEnv:
all_docs = {"index": None}
titles = {
"index": type("TitleNode", (), {"astext": lambda: "Index Page"})()
}
toctree_includes = {"index": []}
manager.set_env(MockEnv())
manager.set_master_doc("index")
# Test the priority behavior
manager.combine_sources(outdir, srcdir)
# Check that output file was created
output_file = f"{outdir}/test.txt"
assert os.path.exists(output_file)
with open(output_file, "r") as f:
content = f.read()
# The system should prefer RST over MD for the "index" docname
# But since both files exist and the second phase adds remaining files,
# both will be included. The test verifies that RST appears first
# (indicating it was found first in the priority order)
assert "RST content for index" in content
# Find positions to verify order
rst_pos = content.find("RST content for index")
md_pos = content.find("Markdown content for index")
# RST should come before MD (due to priority in toctree processing)
assert rst_pos < md_pos, "RST content should appear before MD content"
def test_summary_default_uses_first_paragraph():
"""
Test that summary defaults to first paragraph of root document when not configured.
"""
from docutils import nodes
from docutils.frontend import OptionParser
from docutils.parsers.rst import Parser
from docutils.utils import new_document
from sphinx_llms_txt import build_finished, doctree_resolved
# Create a proper document with settings
settings = OptionParser(components=(Parser,)).get_default_values()
doctree = new_document("<rst-doc>", settings)
title = nodes.title(text="Test Title")
paragraph = nodes.paragraph(
text="This is the first paragraph that should be used as summary."
)
doctree.append(title)
doctree.append(paragraph)
# Mock Sphinx app
class MockApp:
class Config:
master_doc = "index"
llms_txt_summary = None # Not configured
llms_txt_file = True
llms_txt_filename = "llms.txt"
llms_txt_title = None
llms_txt_full_file = True
llms_txt_full_filename = "llms-full.txt"
llms_txt_full_max_size = None
llms_txt_directives = []
llms_txt_exclude = []
html_baseurl = ""
config = Config()
outdir = "/tmp/build"
srcdir = "/tmp/source"
class Env:
titles = {
"index": type("TitleNode", (), {"astext": lambda self: "Test Title"})()
}
env = Env()
app = MockApp()
# Reset the global state
import sphinx_llms_txt
sphinx_llms_txt._root_first_paragraph = ""
# Call doctree_resolved to extract the first paragraph
doctree_resolved(app, doctree, "index")
# Verify the first paragraph was extracted
assert (
sphinx_llms_txt._root_first_paragraph
== "This is the first paragraph that should be used as summary."
)
# Mock the manager methods to avoid actual file operations
original_combine_sources = sphinx_llms_txt._manager.combine_sources
sphinx_llms_txt._manager.combine_sources = lambda outdir, srcdir: None
# Call build_finished and verify the summary is set correctly
build_finished(app, None)
# Check that the summary was properly configured
assert (
sphinx_llms_txt._manager.config["llms_txt_summary"]
== "This is the first paragraph that should be used as summary."
)
# Restore original method
sphinx_llms_txt._manager.combine_sources = original_combine_sources