diff --git a/VERSION b/VERSION index 9e11b32..d15723f 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -0.3.1 +0.3.2 diff --git a/docs/razdeljane-na-sekcii.md b/docs/razdeljane-na-sekcii.md index 731be5e..dd541a5 100644 --- a/docs/razdeljane-na-sekcii.md +++ b/docs/razdeljane-na-sekcii.md @@ -1,6 +1,6 @@ # Разделяне на произволен текст на секции -Този документ описва **реалното** поведение на RIP Help System при разделяне на help-файл на секции. Източникът е `help_processor.py` след последните корекции (вкл. `merge_short_sections`). Не е маркетингово резюме. +Този документ описва **реалното** поведение на RIP Help System при разделяне на help-файл на секции. Източникът е `help_processor.py` след последните корекции (вкл. `merge_short_sections` / `merge_preamble_sections`). Не е маркетингово резюме. Генерирането на ключови думи е **извън обхвата**, освен един ред в края. @@ -18,7 +18,7 @@ python docs/md_to_pdf.py 1. Избира парсер по разширение (`.html` / `.htm`, `.docx`, `.doc`, `.pdf`, `.txt`). 2. Парсерът връща списък от `Section(title, text, level, images, html_text)`. -3. Върху списъка се вика `merge_short_sections()`. +3. Върху списъка се вика `merge_short_sections()`, после `merge_preamble_sections()`. 4. Празните блокове се филтрират при запис; заглавие без тяло се запазва (виж §8). ZIP не е формат за разделяне: разопакова се и всеки вътрешен файл минава по същия път. @@ -29,11 +29,19 @@ ZIP не е формат за разделяне: разопакова се и ## 2. Какво е граница на секция -**Граница = открито заглавие.** Текущата секция се затваря (`flush`) и започва нова. Текстът на заглавието става `title` и **не** се копира в тялото на секцията. +**Граница = открито заглавие ниво 1** (H1 / Heading 1 / Заглавие 1) в HTML и DOCX. Текущата секция се затваря (`flush`) и започва нова. Текстът на заглавието става `title` и **не** се копира в тялото на секцията. -Всичко, което не е заглавие — параграфи, списъци, таблици, картинки — се добавя към **текущата** секция. Таблица никога не отваря нова секция. +В HTML/DOCX **не** отварят нова секция: -Ако преди първото заглавие има съдържание, то образува секция с **празно** `title`. +- Word **Title** / **Наименование** / CSS `Title` / `MsoTitle` — корица; става заглавие на преамбюла (или ред в тялото, ако вече има съдържание); +- „**Съдържание**“ / Contents / TOC и еквиваленти — остават в тялото; +- H2–H6, Subtitle / Подзаглавие, изцяло bold параграфи — влизат като редове в **текущата** секция. + +TXT и PDF запазват своите правила (номерирани редове / по-голям шрифт) — там ниво 2 е нормална граница. + +Всичко, което не е граница — параграфи, списъци, таблици, картинки — се добавя към **текущата** секция. + +Ако преди първото H1 има съдържание (корица, TOC), то образува преамбюл секция (с title от корицата, ако има). Ако парсерът не открие нито една секция, файлът става **една** секция без заглавие, с целия извлечен текст. @@ -49,32 +57,32 @@ ZIP не е формат за разделяне: разопакова се и Разпознават се поне: -| Токен (примери) | Ниво | +| Токен (примери) | Ниво / роля | |---|---| -| Heading 1 / heading1 | 1 | -| Heading 2 / 3 | 2 / 3 | -| Title | 1 | -| Subtitle | 2 | -| **Заглавие** / **Заглавие 1** | 1 | -| Заглавие 2 / 3 | 2 / 3 | -| Подзаглавие | 2 | -| Наименование | 1 | -| Überschrift / Überschrift 1 | 1 | +| Heading 1 / heading1 | 1 — **граница** | +| Heading 2 / 3 | 2 / 3 — в тялото, не граница | +| **Title** / MsoTitle / **Наименование** | корица — не граница | +| Subtitle / Подзаглавие | 2 — в тялото | +| **Заглавие** / **Заглавие 1** | 1 — **граница** | +| Заглавие 2 / 3 | 2 / 3 — в тялото | +| Überschrift / Überschrift 1 | 1 — **граница** | | MsoHeading1 / 2 / 3 | 1 / 2 / 3 | -Число над 3 се ограничава до 3. `Subtitle` / `Подзаглавие` без цифра → ниво 2; останалите именувани стилове без цифра → ниво 1. +Число над 3 се ограничава до 3. `Subtitle` / `Подзаглавие` без цифра → ниво 2; останалите именувани heading-стилове без цифра → ниво 1. **Title** / **Наименование** не връщат ниво за сегментиране. Ако стилът не е heading, има fallback: Word `outlineLvl` 0–2 **и** текст под 120 символа → ниво = `outlineLvl + 1`. -### 3.2. Bold като заглавие +### 3.2. Bold като подзаглавие (не граница) -Ако няма стилово ниво, параграфът се приема за заглавие ниво 2, когато: +Ако няма стилово ниво 1, параграфът се приема за **текст в тялото** (бивш „heading“ ниво 2), когато: - видимият текст е непразен и ≤ 120 символа; - всички run-ове с текст са **bold**; - стилът **не** започва с `list`; - в параграфа **няма** картинки. +Такъв ред **не** отваря нова секция — остава под текущото H1. + ### 3.3. Таблици и картинки Таблица: всеки ред се сплесква до `клетка | клетка | клетка` и редовете се добавят в тялото на текущата секция. @@ -103,9 +111,10 @@ ZIP не е формат за разделяне: разопакова се и ### 5.1. Заглавие -1. Тагове: `h1`→1, `h2`→2, `h3`–`h6`→3. -2. CSS class съдържа `heading`, `заглавие`, `title`, `subtitle`, `msoheading` или `überschrift`, с опционална цифра. Така `MsoHeading1` / `MsoNormal` **не** се бъркат: само heading-класовете режат секция. Пример: Word HTML `

`. -3. `p` или `div` с текст < 120 символа, изцяло обвит в `` / `` → ниво 2. +1. Тагове: `h1`→ граница (ниво 1); `h2`–`h6`→ в тялото (ниво 2/3). +2. CSS class с `heading` / `заглавие` / `msoheading` / `überschrift` (+ опционална цифра). `Title` / `MsoTitle` / `наименование` са **корица**, не граница. `subtitle` / `подзаглавие` → в тялото. +3. `p` или `div` с текст < 120 символа, изцяло обвит в `` / `` → в тялото (не граница). +4. Текст „Съдържание“ / Contents / TOC (дори като `h1`) → в тялото на преамбюла. ### 5.2. Съдържание @@ -142,20 +151,25 @@ ZIP не е формат за разделяне: разопакова се и --- -## 8. Сливане на кратки секции (`merge_short_sections`) +## 8. Сливане на кратки секции (`merge_short_sections` / `merge_preamble_sections`) Константа: `MIN_SECTION_TOKENS = 60` (думи в **тялото**, `str.split()`). -Правило след последните корекции: +`merge_short_sections`: - Секция **с непразно заглавие** **не се слива**, дори тялото да е късо или празно. - Секция **без заглавие** и с под 60 думи се залепва към предишната (текст, картинки, HTML). - Няма предишна секция → късият untitled блок остава сам. +`merge_preamble_sections` (след горното): + +- Ако първата секция е корица с късо/празно тяло, а втората е озаглавена „Съдържание“/TOC → сливат се в един преамбюл (title = корицата; TOC заглавието влиза в тялото). + Следствия: -- Последователни заглавия без тяло („Глава 1“ веднага следвано от „1.1 Увод“) остават **две** секции. -- Къс абзац без заглавие след секция се присъединява към нея, вместо да стане отделен къс запис. +- Корица + съдържание + първа глава (H1) → **две** секции (не четири от Title / TOC / H1 / H2). +- Последователни H1 без тяло („Глава 1“ веднага следвано от друг H1) остават **две** секции. +- H2 / bold подзаглавия („Какво е …?“) не режат секция — остават в тялото на текущото H1. - Untitled текст в началото на файла остава отделна секция (докато не е под прага *и* няма към какво да се слее — първият елемент не се слива). При запис в `process_file()`: diff --git a/docs/razdeljane-na-sekcii.pdf b/docs/razdeljane-na-sekcii.pdf index 5f4479f..993b2b7 100644 Binary files a/docs/razdeljane-na-sekcii.pdf and b/docs/razdeljane-na-sekcii.pdf differ diff --git a/help_processor.py b/help_processor.py index 36fc4e9..2b595eb 100644 --- a/help_processor.py +++ b/help_processor.py @@ -562,15 +562,33 @@ _HEADING_TOKEN_RE = re.compile( r"(\d+)?$", re.I, ) +# Word „Title“ / MsoTitle ≠ Heading 1 — корица, не граница на секция. +_COVER_TITLE_TOKENS = frozenset({ + "title", "mstotitle", "наименование", "doctitle", "documenttitle", +}) _HTML_HEADING_CLASS_RE = re.compile( - r"(?:^|[\s_-])(?:heading|заглавие|title|subtitle|msoheading|überschrift)\s*(\d+)?(?:$|[\s_-])", + r"(?:^|[\s_-])(?:heading|заглавие|msoheading|überschrift)\s*(\d+)?(?:$|[\s_-])", + re.I, +) +_HTML_COVER_CLASS_RE = re.compile( + r"(?:^|[\s_-])(?:mso)?title(?:\d+)?(?:$|[\s_-])|" + r"(?:^|[\s_-])наименование(?:\d+)?(?:$|[\s_-])", + re.I, +) +_HTML_SUBTITLE_CLASS_RE = re.compile( + r"(?:^|[\s_-])(?:subtitle|подзаглавие)(?:\d+)?(?:$|[\s_-])", + re.I, +) +_TOC_HEADING_RE = re.compile( + r"^(съдържание|съдържанието|contents|table of contents|toc|" + r"inhaltsverzeichnis|оглавление|содержание)\s*:?\s*$", re.I, ) _HEADING_LEVEL = { "heading1": 1, "heading2": 2, "heading3": 3, "heading4": 3, "heading5": 3, "heading6": 3, - "title": 1, "subtitle": 2, "msoheading1": 1, "msoheading2": 2, "msoheading3": 3, + "subtitle": 2, "msoheading1": 1, "msoheading2": 2, "msoheading3": 3, "заглавие": 1, "заглавие1": 1, "заглавие2": 2, "заглавие3": 3, - "подзаглавие": 2, "наименование": 1, + "подзаглавие": 2, "überschrift": 1, "überschrift1": 1, "überschrift2": 2, "überschrift3": 3, } @@ -579,10 +597,20 @@ def _compact_style_token(s: str) -> str: return re.sub(r"[\s_\-]+", "", (s or "").strip().lower()) +def _is_cover_title_token(token: str) -> bool: + return _compact_style_token(token) in _COVER_TITLE_TOKENS + + +def _is_toc_heading(text: str) -> bool: + return bool(_TOC_HEADING_RE.match((text or "").strip())) + + def _heading_level_from_token(token: str) -> Optional[int]: t = _compact_style_token(token) if not t: return None + if _is_cover_title_token(t): + return None if t in _HEADING_LEVEL: return _HEADING_LEVEL[t] m = _HEADING_TOKEN_RE.match(t) or _HEADING_TOKEN_RE.match((token or "").strip()) @@ -594,21 +622,38 @@ def _heading_level_from_token(token: str) -> Optional[int]: kind = (m.group(1) or "").lower() if kind in ("subtitle", "подзаглавие"): return 2 + if kind in ("title", "наименование"): + return None return 1 -def _docx_heading_level(para) -> Optional[int]: - """Heading 1 / Заглавие 1 / style_id / outlineLvl — включително локализиран Word.""" +def _docx_style_tokens(para) -> list[str]: style = getattr(para, "style", None) + tokens: list[str] = [] seen: set[int] = set() cur = style while cur is not None and id(cur) not in seen: seen.add(id(cur)) for token in (getattr(cur, "style_id", None), getattr(cur, "name", None)): - lvl = _heading_level_from_token(str(token or "")) - if lvl: - return lvl + if token: + tokens.append(str(token)) cur = getattr(cur, "base_style", None) + return tokens + + +def _docx_is_cover_title(para) -> bool: + return any(_is_cover_title_token(t) for t in _docx_style_tokens(para)) + + +def _docx_heading_level(para) -> Optional[int]: + """Heading 1 / Заглавие 1 / style_id / outlineLvl — включително локализиран Word. + Word Title / Наименование не са граница (виж _docx_is_cover_title).""" + if _docx_is_cover_title(para): + return None + for token in _docx_style_tokens(para): + lvl = _heading_level_from_token(token) + if lvl: + return lvl try: pPr = para._element.pPr if pPr is not None and pPr.outlineLvl is not None: @@ -628,11 +673,28 @@ def _is_bold_heading_text(text: str, runs) -> bool: return bool(useful) and all(bool(r.bold) for r in useful) +def _html_is_cover_title(el) -> bool: + raw = el.get("class") if hasattr(el, "get") else None + classes: list[str] + if isinstance(raw, str): + classes = raw.split() + else: + classes = list(raw or []) + for c in classes: + if _is_cover_title_token(c): + return True + return bool(_HTML_COVER_CLASS_RE.search(" ".join(classes))) + + def _html_heading_level(el) -> Optional[int]: name = (getattr(el, "name", None) or "").lower() if name in _HTML_HEADING_MAP: return _HTML_HEADING_MAP[name] cls = " ".join(el.get("class") or []) if hasattr(el, "get") else "" + if _HTML_COVER_CLASS_RE.search(cls): + return None + if _HTML_SUBTITLE_CLASS_RE.search(cls): + return 2 m = _HTML_HEADING_CLASS_RE.search(cls) if m: n = m.group(1) @@ -703,7 +765,7 @@ def parse_html(path: Path) -> list[Section]: blocks = [] for el in body.find_all(_HTML_BLOCK_TAGS + ["img", "div"]): name = (el.name or "").lower() - if name == "div" and not _html_heading_level(el): + if name == "div" and not _html_heading_level(el) and not _html_is_cover_title(el): continue if any(id(par) in consumed for par in el.parents): continue @@ -725,12 +787,38 @@ def parse_html(path: Path) -> list[Section]: sec.html_text = "\n".join(sec_html) if sec_html else None sections.append(sec) + def append_block_as_body(el): + _swap_imgs_in_block(el, base_dir, sec_images, img_counter) + _strip_attrs(el) + txt = _html_block_plain_text(el) + if txt: + sec_text.append(txt) + try: + sec_html.append(str(el)) + except Exception: + pass + for el in blocks: + txt = el.get_text(" ", strip=True) + is_cover = _html_is_cover_title(el) heading_lvl = _html_heading_level(el) + + # Корица (Word Title / class Title): заглавие на преамбюла, без нова секция + if is_cover and txt: + if not current_title and not sec_text and not sec_html and not sec_images: + current_title = txt + current_level = 1 + else: + append_block_as_body(el) + continue + if heading_lvl: - txt = el.get_text(" ", strip=True) if not txt: continue + # TOC и H2+ остават в тялото — граница само при H1 / Заглавие 1 + if _is_toc_heading(txt) or heading_lvl >= 2: + append_block_as_body(el) + continue flush() current_title = txt current_level = heading_lvl @@ -742,21 +830,13 @@ def parse_html(path: Path) -> list[Section]: _swap_imgs_in_block(el.parent if el.parent and el.parent.name else el, base_dir, sec_images, img_counter) # ако е заменен с placeholder, добавяме като текст - txt = el.get_text(" ", strip=True) if el.name else "" - if txt: - sec_text.append(txt) - sec_html.append(f"

{txt}

") + img_txt = el.get_text(" ", strip=True) if el.name else "" + if img_txt: + sec_text.append(img_txt) + sec_html.append(f"

{img_txt}

") continue - _swap_imgs_in_block(el, base_dir, sec_images, img_counter) - _strip_attrs(el) - txt = _html_block_plain_text(el) - if txt: - sec_text.append(txt) - try: - sec_html.append(str(el)) - except Exception: - pass + append_block_as_body(el) flush() @@ -852,19 +932,55 @@ def parse_docx(path: Path) -> list[Section]: if not text and not para_imgs: continue + is_cover = _docx_is_cover_title(para) level = _docx_heading_level(para) is_bold_heading = ( not level + and not is_cover and _is_bold_heading_text(text, para.runs) and not style_name.startswith("list") and not para_imgs ) - if level or is_bold_heading: + # Корица (Word Title): заглавие на преамбюла, без нова секция + if is_cover and text: + if not current_title and not buf and not sec_images: + current_title = text + current_level = 1 + else: + buf.append(text) + for im in para_imgs: + img_counter[0] += 1 + im.placeholder = f"img_{img_counter[0]:02d}" + sec_images.append(im) + buf.append(f"[IMG: {im.placeholder}]") + continue + + # TOC и H2+/bold остават в тялото — граница само при Heading 1 / Заглавие 1 + if text and _is_toc_heading(text): + buf.append(text) + for im in para_imgs: + img_counter[0] += 1 + im.placeholder = f"img_{img_counter[0]:02d}" + sec_images.append(im) + buf.append(f"[IMG: {im.placeholder}]") + continue + + if (level or 0) >= 2 or is_bold_heading: + if text: + buf.append(text) + for im in para_imgs: + img_counter[0] += 1 + im.placeholder = f"img_{img_counter[0]:02d}" + sec_images.append(im) + buf.append(f"[IMG: {im.placeholder}]") + continue + + if level == 1: flush() buf, sec_images = [], [] current_title = text - current_level = level or 2 + current_level = 1 continue if text: @@ -1268,6 +1384,29 @@ def merge_short_sections(sections: list[Section]) -> list[Section]: return result +def merge_preamble_sections(sections: list[Section]) -> list[Section]: + """Слива корица (Title) + „Съдържание“/TOC в един преамбюл, ако парсерът ги е разделил.""" + if len(sections) < 2: + return sections + first, second = sections[0], sections[1] + first_words = len((first.text or "").split()) + if not (first.title or "").strip(): + return sections + if first_words >= MIN_SECTION_TOKENS: + return sections + if not _is_toc_heading(second.title or ""): + return sections + + body_parts = [p for p in (second.title, second.text) if (p or "").strip()] + if (first.text or "").strip(): + body_parts.append(first.text.strip()) + merged = Section(first.title, "\n".join(body_parts).strip(), first.level) + merged.images = (first.images or []) + (second.images or []) + html_parts = [h for h in (first.html_text, second.html_text) if h] + merged.html_text = "\n".join(html_parts) if html_parts else None + return [merged] + list(sections[2:]) + + def clean_text(text: str) -> str: """Collapse spaces/tabs but keep newlines (lists, paragraphs).""" text = text.replace("\r\n", "\n").replace("\r", "\n") @@ -1504,6 +1643,7 @@ def process_file( return _file_result(rel, file_index, existing_codes, saved=0) sections = merge_short_sections(sections) + sections = merge_preamble_sections(sections) remove_section_outputs( output_dir, diff --git a/tests/fixtures/nespertcam_launcher_preamble.html b/tests/fixtures/nespertcam_launcher_preamble.html new file mode 100644 index 0000000..024c41b --- /dev/null +++ b/tests/fixtures/nespertcam_launcher_preamble.html @@ -0,0 +1,22 @@ + + + + +

NESPERTCAM Launcher - Пълно ръководство на потребителя

+

Съдържание

+
    +
  1. 1. Преглед на приложението
  2. +
  3. 2. Инсталация и настройка
  4. +
  5. 3. Конфигурация и настройки
  6. +
  7. 4. Използване и функции
  8. +
  9. 5. Описание на интерфейса
  10. +
  11. 6. Решаване на проблеми
  12. +
  13. 7. Разширени функции
  14. +
  15. 8. Технически детайли
  16. +
  17. 9. Поддръжка и контакти
  18. +
+

1. Преглед на приложението

+

Какво е NESPERTCAM Launcher?

+

NESPERTCAM Launcher е специализирано Windows приложение, предназначено за лесно и удобно стартиране на NESPERTCAM64.exe и свързаните с него модули от едно място.

+ + diff --git a/tests/test_section_split_preamble.py b/tests/test_section_split_preamble.py new file mode 100644 index 0000000..f63fd75 --- /dev/null +++ b/tests/test_section_split_preamble.py @@ -0,0 +1,80 @@ +"""Корица + Съдържание + глава 1 → 2 секции (не 4).""" +from pathlib import Path + +from help_processor import ( + merge_preamble_sections, + merge_short_sections, + parse_html, + Section, +) + +FIXTURE = Path(__file__).parent / "fixtures" / "nespertcam_launcher_preamble.html" + + +def _titles(sections): + return [s.title for s in sections] + + +def test_nespertcam_preamble_yields_two_sections(): + sections = merge_preamble_sections(merge_short_sections(parse_html(FIXTURE))) + assert len(sections) == 2 + assert sections[0].title == "NESPERTCAM Launcher - Пълно ръководство на потребителя" + assert sections[1].title == "1. Преглед на приложението" + assert "Съдържание" in sections[0].text + assert "1. Преглед на приложението" in sections[0].text + assert "Какво е NESPERTCAM Launcher?" in sections[1].text + assert "NESPERTCAM64.exe" in sections[1].text + + +def test_h2_and_bold_do_not_split_inside_chapter(): + html = Path(__file__).parent / "fixtures" / "_tmp_h2.html" + html.write_text( + """ +

Глава А

+

Подточка

+

Текст под подточката с достатъчно думи за тяло.

+

Още едно bold подзаглавие

+

Още текст в същата секция.

+

Глава Б

+

Тяло на втората глава.

+ """, + encoding="utf-8", + ) + try: + sections = merge_preamble_sections(merge_short_sections(parse_html(html))) + assert _titles(sections) == ["Глава А", "Глава Б"] + assert "Подточка" in sections[0].text + assert "Още едно bold подзаглавие" in sections[0].text + finally: + html.unlink(missing_ok=True) + + +def test_toc_heading_as_h1_merges_into_cover(): + """Ако „Съдържание“ е H1, merge_preamble го залепва към корицата.""" + secs = [ + Section("Ръководство X", "", 1), + Section("Съдържание", "1. Увод\n2. Край", 1), + Section("1. Увод", "Текст на увода.", 1), + ] + merged = merge_preamble_sections(secs) + assert len(merged) == 2 + assert merged[0].title == "Ръководство X" + assert "Съдържание" in merged[0].text + assert merged[1].title == "1. Увод" + + +def test_heading1_chapters_still_split(): + html = Path(__file__).parent / "fixtures" / "_tmp_chapters.html" + html.write_text( + """ +

1. Първа

Аа аа аа.

+

2. Втора

Бб бб бб.

+

3. Трета

Вв вв вв.

+ """, + encoding="utf-8", + ) + try: + sections = merge_preamble_sections(merge_short_sections(parse_html(html))) + assert _titles(sections) == ["1. Първа", "2. Втора", "3. Трета"] + finally: + html.unlink(missing_ok=True)