You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Research: what else the SCJN's SCOW search gives us that we throw away (ámbito, categoría, materia, vigencia, reform extract, legislative process) #203
nota2md.scjn_api talks to the SCJN's SCOW JSON API (the backend of legislacion.scjn.gob.mx/consulta/buscador, issue #172) for exactly one purpose:
find the idOrdenamiento for a catalogue entry, walk its reform table, and pull
the consolidated text of each reform. Everything else the search answers with is
either dropped on the floor or written into a snapshot header and never read
back.
That is a lot of signal. BusquedaFrase returns a fully faceted search index
over ~41 000 instruments, classified by ámbito, categoría, materia and vigencia, and the API exposes seven endpoints we call zero times. This issue
is the survey of what is actually there (measured live against the API on
2026-09-03) and what it could buy us. No implementation is proposed here — the
point is to decide which of these are worth their own issues.
What we consume today
Field
Where it comes from
What we do with it
idOrdenamiento
BusquedaFrase
addressing; written to the snapshot header
ordenamiento
BusquedaFrase
similarity scoring in elige_ordenamiento
categoriaOrdenamiento
BusquedaFrase
mapped to grupo_de_categoria (LEY/CÓDIGO/CONSTITUCIÓN vs REGLAMENTO) for candidate exclusion
scripts/extract_scjn_titles.py --discover is the one place that uses the
filters as filters, paging ambitoF=FEDERAL × categoriaF∈{LEY,CODIGO,CONSTITUCION}
× vigenciaF=VIGENTE to find laws the catalogue is missing.
What the API returns that we never touch
1. Search-hit fields dropped at parse time
ResponseDocumentoSIL (the search-result schema) carries, per hit:
resumen — a one-paragraph human abstract of the instrument, written by
the SCJN. For LEY FEDERAL DEL TRABAJO: "Ley que es de observancia general en
toda la República y rige las relaciones de trabajo comprendidas en el artículo
123, Apartado 'A', de la Constitución." We have no abstract for any law
anywhere in the project.
vigencia — the instrument's own status, and it is not a boolean: the
facet counts show seven distinct values (VIGENTE, NO VIGENTE, ABROGADO (A), SIN EFECTO, DEROGADO (A), INEFICAZ, EXTINGUIDO (A)).
We never record whether a law we crawled is still in force.
reformaExtracto — the text of the decree's operative clause for the
newest reform: "SE REFORMAN LOS ARTICULOS 304; 307; 308, PARRAFO PRIMERO; 309; 310; Y LA DENOMINACION DEL CAPITULO XI DEL TITULO SEXTO; Y SE ADICIONA UN ARTICULO 305 BIS...". The same field appears as reforma on every row of
the Reforma table, i.e. for every reform of every law.
estado, municipio, pais — null for federal instruments, the whole
addressing scheme for the 40 095 state-level ones.
totalResultadosArticulos / articuloOrdenamientos — with consultaArticulos=1 the same call does full-text search inside article
bodies and returns the matching articles with <em> hit highlighting.
2. The facets
Every BusquedaFrase response also carries facetasAmbito, facetasCategoria, facetasVigencia, facetasEstado, facetasMateria, facetasMunicpio (sic) — {categoria, conteo} pairs. Querying q=ley with no filters gives the shape of
the whole index:
materia (25 values): FISCAL 26 534, PRESUPUESTAL 9 854, ADMINISTRATIVO 8 456, ADMINISTRATIVO FISCAL 2 343, CIVIL 441, CONSTITUCIONAL 371, PENAL 343, LABORAL 202, ELECTORAL 176, PUEBLOS INDÍGENAS Y AFROMEXICANOS 145, SEGURIDAD SOCIAL 68, …
plus compound values (ADMINISTRATIVA Y CIVIL, PENAL Y DE TRABAJO)
estado: all 32 entities, OAXACA 7 435 down to BAJA CALIFORNIA SUR 343
The facets are a free, no-extra-request census. They are also the only way to
enumerate the legal vocabularies — there is no catalogue endpoint for them.
3. Seven unused endpoints
From SCOW-API/swagger/v1/swagger.json, we call three of thirteen paths:
Endpoint
What it does
Verified
GET /api/SCOW/ProcesosLegislativos
the legislative history of one reform: exposición de motivos, dictamen de origen, discusión, dictamen revisora, … each with a PDF and a tipoProceso
yes — 48 of the LFT's 61 reforms carry tieneProcesos: true, and the call returns the whole ordered chain with Gaceta numbers and dates
GET /api/SCOW/ArticulosOrdenamiento
full-text search restricted to one law, returning the matching articles with hit highlighting
yes — idOrdenamiento=410&fraseBusqueda=teletrabajo returns CAPÍTULO XII BIS and art. 330-A
POST /api/SCOWDocs/CronologicaArticulo
the chronology of a single article: every version it has had across reforms
yes (returns articulos with fechaPublicacion per version)
POST /api/SCOW/BusquedaAvanzada
exact-phrase, excluded-terms, date-range and sort-order search over the same index
not exercised
GET /api/SCOW/Articulo
one article by (ordenamientoId, reformaId, articuloId)
not exercised
POST /api/SCOWDocs/DocumentosReforma / DocumentosArticulos
reform/article documents
not exercised
GET /api/SCOW/Autocompletado
title autocomplete
returns empty for the bodies tried; may need a different tipoPublicacion
tipoPublicacion looks inert: 1, 2 and 3 answer identically (same totals, same
facets) — worth confirming before anyone builds on it.
4. Per-article fields we drop
ArticuloOrdenamiento carries id (a stable article id), articuloVersion
(a version counter — art. 1 of the LFT is at version 2), fechaActualizacion
(when the SCJN last touched that article: 02/07/2019 16:29:19) and per-article vigencia. We keep only numero, orden, referencia, contenido.
What this could buy us
Roughly in order of how much I think it is worth:
A vigencia field on the corpus. We crawl and republish laws with no
record of whether they are still in force. scjn-leyes currently cannot
answer "give me the laws in force today" without going back to the API. This
is a one-field change to cabecera and the index, and it is the cheapest
real gap on this list.
materia as a usable label, not a header comment. We already write it
and never read it. Promoting it to the release index turns scjn-leyes into
a classified corpus — the natural basis for topic-conditioned evaluation of
anything downstream (retrieval, classification, summarization), and a
ready-made stratification variable for sampling. The compound values
(PENAL Y ADMINISTRATIVA) need a decision: multi-label or verbatim.
reformaExtracto / reforma as ground truth for the reconstruction.
This is the one I would prioritize for correctness work. Every reform row
states, in the decree's own words, which articles it reformed, added or
derogated. reconstruct_legal_provisions replays reforms onto a base text
and has no independent check that it touched the right articles; parsing
these clauses (SE REFORMAN LOS ARTICULOS 304; 307; …) gives an article-level
expectation to diff the replay against — the same kind of validation issue Fase 3 — Replace texto_vigente's ground truth for reconstruct_legal_provisions() #188 did with the SCJN's own text, but per article instead of per snapshot.
It also gives md2akn a real refersTo target set.
The legislative process corpus.ProcesosLegislativos is a second corpus
the project does not have at all: initiative → dictamen → floor discussion,
linked to the reform it produced and therefore, through our existing codNota link, to the DOF publication. For a project about analyzing legal
texts, having the intent documents alongside the enacted text is a
qualitatively new capability (legislative intent, argument mining, before/after
pairs). Cost: one extra request per reform, plus PDFs behind pdf/ruta whose retrieval needs its own look — several are "solicítelo por
correo" placeholders.
resumen as a free abstract. One human-written paragraph per instrument.
Useful as a description in the release index and as reference summaries for
evaluating generated ones.
The facets as a coverage instrument.--discover already pages the
filters; reading the facet counts back would tell us directly how many
federal VIGENTELEYs the SCJN believes exist versus how many the
catalogue has — a coverage number we currently estimate rather than read.
ambito=ESTATAL — the 40 095 instruments we ignore. The project is
federal by design (the DOF is federal), so this is deliberate. But it is
worth stating explicitly that the same crawler, unchanged, reaches all 32
states' legislation, indexed by estado/municipio, and deciding whether
that is out of scope forever or just for now. Note there is no DOF codNota
to link state instruments to — the whole provenance story would differ.
Open questions
Is materia assigned per instrument or per reform? The search hit carries it;
the reform rows do not.
How stable are these vocabularies? Neither categoriaOrdenamiento nor materia has a documented enumeration — the facets are the only listing, and
they are query-dependent.
Does vigencia change retroactively on old snapshots we already published, and
if so does the snapshot header record the value at crawl time or the current one?
Same posture caveat as always: the API is public but has no stability
contract, and fuente: scjn still does not make the SCJN an official source.
Anything adopted from here is metadata about the law, and the DOF/SIDOF
remains the authority on its text.
Suggested next step
Split (1), (2) and (3) into implementation issues — they are small, they touch
code we already own, and (3) is a correctness win rather than a feature. Treat
(4) as its own epic; treat (8) as a scope decision to record, not to build.
Why
nota2md.scjn_apitalks to the SCJN's SCOW JSON API (the backend oflegislacion.scjn.gob.mx/consulta/buscador, issue #172) for exactly one purpose:find the
idOrdenamientofor a catalogue entry, walk its reform table, and pullthe consolidated text of each reform. Everything else the search answers with is
either dropped on the floor or written into a snapshot header and never read
back.
That is a lot of signal.
BusquedaFrasereturns a fully faceted search indexover ~41 000 instruments, classified by ámbito, categoría, materia and
vigencia, and the API exposes seven endpoints we call zero times. This issue
is the survey of what is actually there (measured live against the API on
2026-09-03) and what it could buy us. No implementation is proposed here — the
point is to decide which of these are worth their own issues.
What we consume today
idOrdenamientoBusquedaFraseordenamientoBusquedaFraseelige_ordenamientocategoriaOrdenamientoBusquedaFrasegrupo_de_categoria(LEY/CÓDIGO/CONSTITUCIÓN vs REGLAMENTO) for candidate exclusionambitoBusquedaFraseFEDERALmateriaBusquedaFraseiweightBusquedaFrasefechaPublicadoBusquedaFraseReformacategoriaReforma,seccionPublicacion,pdf,tieneArticulosArticulosscripts/extract_scjn_titles.py --discoveris the one place that uses thefilters as filters, paging
ambitoF=FEDERAL×categoriaF∈{LEY,CODIGO,CONSTITUCION}×
vigenciaF=VIGENTEto find laws the catalogue is missing.What the API returns that we never touch
1. Search-hit fields dropped at parse time
ResponseDocumentoSIL(the search-result schema) carries, per hit:resumen— a one-paragraph human abstract of the instrument, written bythe SCJN. For
LEY FEDERAL DEL TRABAJO: "Ley que es de observancia general entoda la República y rige las relaciones de trabajo comprendidas en el artículo
123, Apartado 'A', de la Constitución." We have no abstract for any law
anywhere in the project.
vigencia— the instrument's own status, and it is not a boolean: thefacet counts show seven distinct values (
VIGENTE,NO VIGENTE,ABROGADO (A),SIN EFECTO,DEROGADO (A),INEFICAZ,EXTINGUIDO (A)).We never record whether a law we crawled is still in force.
reformaExtracto— the text of the decree's operative clause for thenewest reform:
"SE REFORMAN LOS ARTICULOS 304; 307; 308, PARRAFO PRIMERO; 309; 310; Y LA DENOMINACION DEL CAPITULO XI DEL TITULO SEXTO; Y SE ADICIONA UN ARTICULO 305 BIS...". The same field appears asreformaon every row ofthe
Reformatable, i.e. for every reform of every law.estado,municipio,pais— null for federal instruments, the wholeaddressing scheme for the 40 095 state-level ones.
totalResultadosArticulos/articuloOrdenamientos— withconsultaArticulos=1the same call does full-text search inside articlebodies and returns the matching articles with
<em>hit highlighting.2. The facets
Every
BusquedaFraseresponse also carriesfacetasAmbito,facetasCategoria,facetasVigencia,facetasEstado,facetasMateria,facetasMunicpio(sic) —{categoria, conteo}pairs. Queryingq=leywith no filters gives the shape ofthe whole index:
ESTATAL40 095,FEDERAL1 061LEY38 440,ACUERDO (S)347,LINEAMIENTOS63,DECRETO50,DECLARATORIA18,CONTRATO14,CODIGO2, …FISCAL26 534,PRESUPUESTAL9 854,ADMINISTRATIVO8 456,ADMINISTRATIVO FISCAL2 343,CIVIL441,CONSTITUCIONAL371,PENAL343,LABORAL202,ELECTORAL176,PUEBLOS INDÍGENAS Y AFROMEXICANOS145,SEGURIDAD SOCIAL68, …plus compound values (
ADMINISTRATIVA Y CIVIL,PENAL Y DE TRABAJO)OAXACA7 435 down toBAJA CALIFORNIA SUR343The facets are a free, no-extra-request census. They are also the only way to
enumerate the legal vocabularies — there is no catalogue endpoint for them.
3. Seven unused endpoints
From
SCOW-API/swagger/v1/swagger.json, we call three of thirteen paths:GET /api/SCOW/ProcesosLegislativostipoProcesotieneProcesos: true, and the call returns the whole ordered chain with Gaceta numbers and datesGET /api/SCOW/ArticulosOrdenamientoidOrdenamiento=410&fraseBusqueda=teletrabajoreturns CAPÍTULO XII BIS and art. 330-APOST /api/SCOWDocs/CronologicaArticuloarticuloswithfechaPublicacionper version)POST /api/SCOW/BusquedaAvanzadaGET /api/SCOW/Articulo(ordenamientoId, reformaId, articuloId)POST /api/SCOWDocs/DocumentosReforma/DocumentosArticulosGET /api/SCOW/AutocompletadotipoPublicaciontipoPublicacionlooks inert: 1, 2 and 3 answer identically (same totals, samefacets) — worth confirming before anyone builds on it.
4. Per-article fields we drop
ArticuloOrdenamientocarriesid(a stable article id),articuloVersion(a version counter — art. 1 of the LFT is at version 2),
fechaActualizacion(when the SCJN last touched that article:
02/07/2019 16:29:19) and per-articlevigencia. We keep onlynumero,orden,referencia,contenido.What this could buy us
Roughly in order of how much I think it is worth:
vigenciafield on the corpus. We crawl and republish laws with norecord of whether they are still in force.
scjn-leyescurrently cannotanswer "give me the laws in force today" without going back to the API. This
is a one-field change to
cabeceraand the index, and it is the cheapestreal gap on this list.
materiaas a usable label, not a header comment. We already write itand never read it. Promoting it to the release index turns
scjn-leyesintoa classified corpus — the natural basis for topic-conditioned evaluation of
anything downstream (retrieval, classification, summarization), and a
ready-made stratification variable for sampling. The compound values
(
PENAL Y ADMINISTRATIVA) need a decision: multi-label or verbatim.reformaExtracto/reformaas ground truth for the reconstruction.This is the one I would prioritize for correctness work. Every reform row
states, in the decree's own words, which articles it reformed, added or
derogated.
reconstruct_legal_provisionsreplays reforms onto a base textand has no independent check that it touched the right articles; parsing
these clauses (
SE REFORMAN LOS ARTICULOS 304; 307; …) gives an article-levelexpectation to diff the replay against — the same kind of validation issue
Fase 3 — Replace texto_vigente's ground truth for reconstruct_legal_provisions() #188 did with the SCJN's own text, but per article instead of per snapshot.
It also gives
md2akna realrefersTotarget set.ProcesosLegislativosis a second corpusthe project does not have at all: initiative → dictamen → floor discussion,
linked to the reform it produced and therefore, through our existing
codNotalink, to the DOF publication. For a project about analyzing legaltexts, having the intent documents alongside the enacted text is a
qualitatively new capability (legislative intent, argument mining, before/after
pairs). Cost: one extra request per reform, plus PDFs behind
pdf/rutawhose retrieval needs its own look — several are "solicítelo porcorreo" placeholders.
resumenas a free abstract. One human-written paragraph per instrument.Useful as a description in the release index and as reference summaries for
evaluating generated ones.
fechaActualizacion+articuloVersionfor incremental crawling. Todayfreshness is decided per instrument (
actualizadoincatalogo.json,issue SCJN-leyes: cobertura completa del catálogo de Diputados y registrar la consulta de búsqueda en la cabecera #124/Fase 1 — Build catalogo.json without Diputados: instrument discovery and
actualizado#186). Per-article timestamps would let a refresh re-fetch onlywhat changed, and would make "which articles changed between these two
snapshots" answerable without diffing text.
--discoveralready pages thefilters; reading the facet counts back would tell us directly how many
federal
VIGENTELEYs the SCJN believes exist versus how many thecatalogue has — a coverage number we currently estimate rather than read.
ambito=ESTATAL— the 40 095 instruments we ignore. The project isfederal by design (the DOF is federal), so this is deliberate. But it is
worth stating explicitly that the same crawler, unchanged, reaches all 32
states' legislation, indexed by
estado/municipio, and deciding whetherthat is out of scope forever or just for now. Note there is no DOF
codNotato link state instruments to — the whole provenance story would differ.
Open questions
materiaassigned per instrument or per reform? The search hit carries it;the reform rows do not.
categoriaOrdenamientonormateriahas a documented enumeration — the facets are the only listing, andthey are query-dependent.
vigenciachange retroactively on old snapshots we already published, andif so does the snapshot header record the value at crawl time or the current one?
tipoPublicacionappears to be ignored by the backend — confirm, and find outwhether treaties (issue Fase 4 — Delete leyesmx, the Diputados code, and the collection abstraction #189 dropped
leyesmx's treaty pairing) are reachablethrough some other parameter.
contract, and
fuente: scjnstill does not make the SCJN an official source.Anything adopted from here is metadata about the law, and the DOF/SIDOF
remains the authority on its text.
Suggested next step
Split (1), (2) and (3) into implementation issues — they are small, they touch
code we already own, and (3) is a correctness win rather than a feature. Treat
(4) as its own epic; treat (8) as a scope decision to record, not to build.