Skip to content

Research: what else the SCJN's SCOW search gives us that we throw away (ámbito, categoría, materia, vigencia, reform extract, legislative process) #203

Description

@mgraffg

Why

nota2md.scjn_api talks to the SCJN's SCOW JSON API (the backend of
legislacion.scjn.gob.mx/consulta/buscador, issue #172) for exactly one purpose:
find the idOrdenamiento for a catalogue entry, walk its reform table, and pull
the consolidated text of each reform. Everything else the search answers with is
either dropped on the floor or written into a snapshot header and never read
back.

That is a lot of signal. BusquedaFrase returns a fully faceted search index
over ~41 000 instruments, classified by ámbito, categoría, materia and
vigencia, and the API exposes seven endpoints we call zero times. This issue
is the survey of what is actually there (measured live against the API on
2026-09-03) and what it could buy us. No implementation is proposed here — the
point is to decide which of these are worth their own issues.

What we consume today

Field Where it comes from What we do with it
idOrdenamiento BusquedaFrase addressing; written to the snapshot header
ordenamiento BusquedaFrase similarity scoring in elige_ordenamiento
categoriaOrdenamiento BusquedaFrase mapped to grupo_de_categoria (LEY/CÓDIGO/CONSTITUCIÓN vs REGLAMENTO) for candidate exclusion
ambito BusquedaFrase one tiebreak: prefer FEDERAL
materia BusquedaFrase written to the header, never read
iweight BusquedaFrase carried on the dataclass, unused
fechaPublicado BusquedaFrase carried, unused
reform table rows Reforma dates, categoriaReforma, seccionPublicacion, pdf, tieneArticulos
article rows Articulos the Markdown body

scripts/extract_scjn_titles.py --discover is the one place that uses the
filters as filters, paging ambitoF=FEDERAL × categoriaF∈{LEY,CODIGO,CONSTITUCION}
× vigenciaF=VIGENTE to find laws the catalogue is missing.

What the API returns that we never touch

1. Search-hit fields dropped at parse time

ResponseDocumentoSIL (the search-result schema) carries, per hit:

  • resumen — a one-paragraph human abstract of the instrument, written by
    the SCJN. For LEY FEDERAL DEL TRABAJO: "Ley que es de observancia general en
    toda la República y rige las relaciones de trabajo comprendidas en el artículo
    123, Apartado 'A', de la Constitución."
    We have no abstract for any law
    anywhere in the project.
  • vigencia — the instrument's own status, and it is not a boolean: the
    facet counts show seven distinct values (VIGENTE, NO VIGENTE,
    ABROGADO (A), SIN EFECTO, DEROGADO (A), INEFICAZ, EXTINGUIDO (A)).
    We never record whether a law we crawled is still in force.
  • reformaExtracto — the text of the decree's operative clause for the
    newest reform: "SE REFORMAN LOS ARTICULOS 304; 307; 308, PARRAFO PRIMERO; 309; 310; Y LA DENOMINACION DEL CAPITULO XI DEL TITULO SEXTO; Y SE ADICIONA UN ARTICULO 305 BIS...". The same field appears as reforma on every row of
    the Reforma table, i.e. for every reform of every law.
  • estado, municipio, pais — null for federal instruments, the whole
    addressing scheme for the 40 095 state-level ones.
  • totalResultadosArticulos / articuloOrdenamientos — with
    consultaArticulos=1 the same call does full-text search inside article
    bodies
    and returns the matching articles with <em> hit highlighting.

2. The facets

Every BusquedaFrase response also carries facetasAmbito, facetasCategoria,
facetasVigencia, facetasEstado, facetasMateria, facetasMunicpio (sic) —
{categoria, conteo} pairs. Querying q=ley with no filters gives the shape of
the whole index:

  • ámbito: ESTATAL 40 095, FEDERAL 1 061
  • categoría (20 values): LEY 38 440, ACUERDO (S) 347, LINEAMIENTOS 63,
    DECRETO 50, DECLARATORIA 18, CONTRATO 14, CODIGO 2, …
  • materia (25 values): FISCAL 26 534, PRESUPUESTAL 9 854,
    ADMINISTRATIVO 8 456, ADMINISTRATIVO FISCAL 2 343, CIVIL 441,
    CONSTITUCIONAL 371, PENAL 343, LABORAL 202, ELECTORAL 176,
    PUEBLOS INDÍGENAS Y AFROMEXICANOS 145, SEGURIDAD SOCIAL 68, …
    plus compound values (ADMINISTRATIVA Y CIVIL, PENAL Y DE TRABAJO)
  • estado: all 32 entities, OAXACA 7 435 down to BAJA CALIFORNIA SUR 343

The facets are a free, no-extra-request census. They are also the only way to
enumerate the legal vocabularies — there is no catalogue endpoint for them.

3. Seven unused endpoints

From SCOW-API/swagger/v1/swagger.json, we call three of thirteen paths:

Endpoint What it does Verified
GET /api/SCOW/ProcesosLegislativos the legislative history of one reform: exposición de motivos, dictamen de origen, discusión, dictamen revisora, … each with a PDF and a tipoProceso yes — 48 of the LFT's 61 reforms carry tieneProcesos: true, and the call returns the whole ordered chain with Gaceta numbers and dates
GET /api/SCOW/ArticulosOrdenamiento full-text search restricted to one law, returning the matching articles with hit highlighting yes — idOrdenamiento=410&fraseBusqueda=teletrabajo returns CAPÍTULO XII BIS and art. 330-A
POST /api/SCOWDocs/CronologicaArticulo the chronology of a single article: every version it has had across reforms yes (returns articulos with fechaPublicacion per version)
POST /api/SCOW/BusquedaAvanzada exact-phrase, excluded-terms, date-range and sort-order search over the same index not exercised
GET /api/SCOW/Articulo one article by (ordenamientoId, reformaId, articuloId) not exercised
POST /api/SCOWDocs/DocumentosReforma / DocumentosArticulos reform/article documents not exercised
GET /api/SCOW/Autocompletado title autocomplete returns empty for the bodies tried; may need a different tipoPublicacion

tipoPublicacion looks inert: 1, 2 and 3 answer identically (same totals, same
facets) — worth confirming before anyone builds on it.

4. Per-article fields we drop

ArticuloOrdenamiento carries id (a stable article id), articuloVersion
(a version counter — art. 1 of the LFT is at version 2), fechaActualizacion
(when the SCJN last touched that article: 02/07/2019 16:29:19) and per-article
vigencia. We keep only numero, orden, referencia, contenido.

What this could buy us

Roughly in order of how much I think it is worth:

  1. A vigencia field on the corpus. We crawl and republish laws with no
    record of whether they are still in force. scjn-leyes currently cannot
    answer "give me the laws in force today" without going back to the API. This
    is a one-field change to cabecera and the index, and it is the cheapest
    real gap on this list.
  2. materia as a usable label, not a header comment. We already write it
    and never read it. Promoting it to the release index turns scjn-leyes into
    a classified corpus — the natural basis for topic-conditioned evaluation of
    anything downstream (retrieval, classification, summarization), and a
    ready-made stratification variable for sampling. The compound values
    (PENAL Y ADMINISTRATIVA) need a decision: multi-label or verbatim.
  3. reformaExtracto / reforma as ground truth for the reconstruction.
    This is the one I would prioritize for correctness work. Every reform row
    states, in the decree's own words, which articles it reformed, added or
    derogated
    . reconstruct_legal_provisions replays reforms onto a base text
    and has no independent check that it touched the right articles; parsing
    these clauses (SE REFORMAN LOS ARTICULOS 304; 307; …) gives an article-level
    expectation to diff the replay against — the same kind of validation issue
    Fase 3 — Replace texto_vigente's ground truth for reconstruct_legal_provisions() #188 did with the SCJN's own text, but per article instead of per snapshot.
    It also gives md2akn a real refersTo target set.
  4. The legislative process corpus. ProcesosLegislativos is a second corpus
    the project does not have at all: initiative → dictamen → floor discussion,
    linked to the reform it produced and therefore, through our existing
    codNota link, to the DOF publication. For a project about analyzing legal
    texts, having the intent documents alongside the enacted text is a
    qualitatively new capability (legislative intent, argument mining, before/after
    pairs). Cost: one extra request per reform, plus PDFs behind
    pdf/ruta whose retrieval needs its own look — several are "solicítelo por
    correo" placeholders.
  5. resumen as a free abstract. One human-written paragraph per instrument.
    Useful as a description in the release index and as reference summaries for
    evaluating generated ones.
  6. fechaActualizacion + articuloVersion for incremental crawling. Today
    freshness is decided per instrument (actualizado in catalogo.json,
    issue SCJN-leyes: cobertura completa del catálogo de Diputados y registrar la consulta de búsqueda en la cabecera #124/Fase 1 — Build catalogo.json without Diputados: instrument discovery and actualizado #186). Per-article timestamps would let a refresh re-fetch only
    what changed, and would make "which articles changed between these two
    snapshots" answerable without diffing text.
  7. The facets as a coverage instrument. --discover already pages the
    filters; reading the facet counts back would tell us directly how many
    federal VIGENTE LEYs the SCJN believes exist versus how many the
    catalogue has — a coverage number we currently estimate rather than read.
  8. ambito=ESTATAL — the 40 095 instruments we ignore. The project is
    federal by design (the DOF is federal), so this is deliberate. But it is
    worth stating explicitly that the same crawler, unchanged, reaches all 32
    states' legislation, indexed by estado/municipio, and deciding whether
    that is out of scope forever or just for now. Note there is no DOF codNota
    to link state instruments to — the whole provenance story would differ.

Open questions

  • Is materia assigned per instrument or per reform? The search hit carries it;
    the reform rows do not.
  • How stable are these vocabularies? Neither categoriaOrdenamiento nor
    materia has a documented enumeration — the facets are the only listing, and
    they are query-dependent.
  • Does vigencia change retroactively on old snapshots we already published, and
    if so does the snapshot header record the value at crawl time or the current one?
  • tipoPublicacion appears to be ignored by the backend — confirm, and find out
    whether treaties (issue Fase 4 — Delete leyesmx, the Diputados code, and the collection abstraction #189 dropped leyesmx's treaty pairing) are reachable
    through some other parameter.
  • Same posture caveat as always: the API is public but has no stability
    contract, and fuente: scjn still does not make the SCJN an official source.
    Anything adopted from here is metadata about the law, and the DOF/SIDOF
    remains the authority on its text.

Suggested next step

Split (1), (2) and (3) into implementation issues — they are small, they touch
code we already own, and (3) is a correctness win rather than a feature. Treat
(4) as its own epic; treat (8) as a scope decision to record, not to build.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    researchOpen-ended exploration, not yet a concrete implementation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions