Skip to content

Commit 7cd94df

Browse files
Add Oracle SQL/PGQ support to Awesome-Text2GQL (#68)
* feat: Add OracleDB support and related tests - Updated pyproject.toml to include oracledb dependency. - Introduced new test suite for translating Cypher queries to Oracle SQL/PGQ. - Implemented dataset preparation tests for Oracle integration. - Added live tests for OracleDB client functionality. - Created query generalizer and template instantiator for Oracle SQL/PGQ. - Enhanced corpus combiner to handle Oracle-specific queries and validation. - Included schema parser for generating Oracle DDL statements. * feat: Enhance Oracle SQL PGQ Translator with primary key mapping and strict validation - Added support for node and edge primary key mappings in OracleSqlPgqQueryTranslator. - Introduced strict property validation to ensure properties are defined for variables. - Updated methods to normalize label maps and handle aggregate functions in WITH clauses. - Enhanced validation for translated queries, including handling of string predicates and label predicates. - Improved error handling for missing properties when strict validation is enabled. - Added new command-line arguments for validation timeout and fetch limit in dataset preparation. - Updated tests to cover new features, including primary key mapping and strict validation scenarios. * feat: Add dataset preparation and Oracle vs Neo4j comparison utilities * Add tests for unsupported query features and failure analysis - Enhance `test_detect_unsupported_oracle_sqlpgq_features` with additional assertions for various unsupported query patterns. - Introduce `test_failure_analysis_groups_unsupported_query_shapes` to analyze failure signatures for unsupported queries. - Implement `test_failure_analysis_uses_manifest_for_invalid_schema` to validate schema direction and property checks against a manifest. - Add normalization tests in `test_compare_normalizes_temporal_strings_and_numeric_precision` and `test_compare_normalizes_oracle_and_neo4j_node_identity`. - Create tests for path normalization in `test_compare_normalizes_single_neo4j_path_to_flat_element_sequence`. - Include checks for nondeterministic limits in `test_compare_detects_nondeterministic_limit_without_order_by`. - Expose file stem label aliases in `test_loader_exposes_file_stem_label_aliases`. * Enhance optional match handling and support for correlated optional matches - Introduced `is_supported_correlated_optional_match` to validate correlated optional matches in Cypher queries. - Updated `detect_unsupported_features` to remove "optional_match" feature if correlated optional matches are supported. - Removed redundant optional match translation logic from `cypher2oracle_sqlpgq`. - Added comprehensive tests for various optional match scenarios, including correlated optional matches and their translations to SQL. - Improved handling of optional match clauses in the dataset preparation and query translation processes. * feat: add CypherSchema class for schema validation and property management - Implemented CypherSchema to manage and validate graph schema based on provided configuration. - Added methods for detecting validation issues in Cypher queries, including node and edge label checks, property validation, and unsafe numeric conversions. - Introduced utility functions for parsing Cypher variable labels, property references, and edge relationships. - Included comprehensive handling of schema name aliases and property types. - Ensured deduplication of validation issues for cleaner output. * feat: enhance Oracle SQL PGQ Translator with stage expression correlation and numeric tolerance checks * Enhance Cypher to Oracle SQL/PGQ translation and validation - Introduced checks for unique schema ownership of properties in CypherSchema. - Added detection for unsafe temporal arithmetic in aggregate queries. - Improved handling of broad bounded variable length relationships in translation. - Updated tests to cover new features and edge cases, including disambiguation of complex aggregate property aliases. - Refactored unsupported feature detection to exclude expensive variable length paths. - Enhanced query translation to preserve real ID properties over pseudo identities. - Added stable tiebreakers for ordered queries with limits in comparison functions. * feat: enhance detection of unsupported features with new patterns and tests * feat: add exporter for validated Oracle SQL/PGQ dataset and enhance README * feat: update README and dataset preparation documentation for Oracle SQL/PGQ support
1 parent d02cce8 commit 7cd94df

57 files changed

Lines changed: 21173 additions & 69 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.gitignore‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,13 @@
11
# specific
22
output/
33
corpus/
4+
.history/
5+
dataset/
6+
broken_db/
7+
examples/Oracle_SQLPGQ_Instance/
8+
examples/generated_corpus/oracle_sqlpgq_*.json
9+
examples/generated_corpus/cypher_to_oracle_sqlpgq*.json
10+
test_oracle_sqlpgq_query.json
411

512
# Byte-compiled / optimized / DLL files
613
__pycache__/
@@ -168,4 +175,4 @@ cython_debug/
168175
#.idea/
169176

170177
# poetry
171-
poetry.lock
178+
poetry.lock

‎README.md‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -37,7 +37,7 @@ Awesome-Text2GQL is an AI-assisted framework for Text2GQL dataset construction.
3737

3838
### Generated Benchmark Dataset
3939

40-
The [Text2GQL-Bench](https://arxiv.org/abs/2602.11745)'s dataset is generated by Awesome-Text2GQL framework. It contains 178,184 (Question, Query) pairs spanning 13 domains. The dataset is available at [Text2GQL-Bench_dataset](https://tugraph-web.oss-cn-beijing.aliyuncs.com/tugraph/datasets/text2gql/Text2GraphQueryBenchmark/Text2GQL-Bench_dataset.zip). To run Text2GQL test, please refer to our [Text2GraphQuery-Driver](https://github.com/TuGraph-family/text2graphquery-driver/tree/main).
40+
The [Text2GQL-Bench](https://arxiv.org/abs/2602.11745)'s dataset is generated by Awesome-Text2GQL framework. It contains 178,184 (Question, Query) pairs spanning 13 domains. The dataset is available at [Text2GQL-Bench_dataset](https://tugraph-web.oss-cn-beijing.aliyuncs.com/tugraph/datasets/text2gql/Text2GraphQueryBenchmark/Text2GQL-Bench_dataset.zip). The dataset including the Oracle SQL/PGQ translated queries is available at [Dataset-with-SQL/PGQ](https://objectstorage.us-ashburn-1.oraclecloud.com/p/8dIkuVGsfnRQlP3ifxVDjQP0pmidpadEY18ltEbkPC4PrZyLTxjJdqDjbtWIEYUW/n/ogcs/b/Text2GQL-Bench_dataset/o/Text2GQL-Bench_dataset.zip), it includes 19633 out of 22407 existing queries. To run Text2GQL test, please refer to our [Text2GraphQuery-Driver](https://github.com/TuGraph-family/text2graphquery-driver/tree/main).
4141

4242
## Demo: TuGraph-DB ChatBot
4343

@@ -195,6 +195,15 @@ After all, run:
195195

196196
When the script finishes, the generated corpus will be saved to examples/generated_corpus/{graph_name}_template_corpus.json.
197197

198+
#### Oracle SQL Property Graphs (SQL/PGQ)
199+
200+
Awesome-Text2GQL includes Oracle SQL/PGQ support for schema conversion, graph setup, query translation, corpus generation, validation, and benchmark dataset preparation.
201+
202+
For detailed workflows, see:
203+
204+
- [Oracle SQL/PGQ data generation workflow](./doc/en-us/development/oracle_sqlpgq_data_generation_workflow.md): convert framework/TuGraph-style schemas into Oracle SQL/PGQ artifacts, create local Oracle property graphs, generate deterministic and LLM-based corpora, validate generated queries, and combine corpus outputs.
205+
- [Dataset preparation utilities](./dataset_prep/README.md): translate benchmark Cypher/GQL-like records to Oracle SQL/PGQ, optionally validate them against Oracle, analyze failures, compare Oracle SQL/PGQ results with Neo4j, and export validated datasets.
206+
198207
#### Cypher2GQL
199208

200209
`python ./examples/cypher2gql.py`

‎app/core/clauses/match_clause.py‎

Lines changed: 6 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -24,14 +24,16 @@ class EdgePattern:
2424
class PathPattern:
2525
node_pattern_list: List[NodePattern]
2626
edge_pattern_list: List[EdgePattern]
27+
path_variable: str = ""
2728

2829

2930
class MatchClause(Clause):
30-
def __init__(self, path_pattern: PathPattern):
31+
def __init__(self, path_pattern: PathPattern, optional: bool = False):
3132
self.path_pattern = path_pattern
33+
self.optional = optional
3234

3335
def to_string(self) -> str:
34-
match_string = "MATCH "
36+
match_string = "OPTIONAL MATCH " if self.optional else "MATCH "
3537
path_degree = len(self.path_pattern.edge_pattern_list)
3638
# add first node
3739
node_pattern = self.path_pattern.node_pattern_list[0]
@@ -51,7 +53,7 @@ def to_string(self) -> str:
5153
return match_string
5254

5355
def to_string_cypher(self) -> str:
54-
match_string = "MATCH "
56+
match_string = "OPTIONAL MATCH " if self.optional else "MATCH "
5557
path_degree = len(self.path_pattern.edge_pattern_list)
5658
# add first node
5759
node_pattern = self.path_pattern.node_pattern_list[0]
@@ -82,7 +84,7 @@ def to_string_cypher(self) -> str:
8284
return match_string
8385

8486
def to_string_gql(self) -> str:
85-
match_string = "MATCH "
87+
match_string = "OPTIONAL MATCH " if self.optional else "MATCH "
8688
path_degree = len(self.path_pattern.edge_pattern_list)
8789
# add first node
8890
node_pattern = self.path_pattern.node_pattern_list[0]

‎app/core/clauses/return_clause.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ class ReturnItem:
1010
property: str
1111
alias: str
1212
function_name: str = ""
13+
expression: str = ""
1314

1415

1516
@dataclass
@@ -18,6 +19,7 @@ class SortItem:
1819
property: str
1920
order: str
2021
function_name: str = ""
22+
expression: str = ""
2123

2224

2325
@dataclass

‎app/core/clauses/where_clause.py‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ class CompareExpression:
1010
property: tuple[str, Dict]
1111
comparison_type: str
1212
comparison_value: str
13+
raw_expression: str = ""
1314

1415

1516
class WhereClause(Clause):

‎app/core/generalizer/query_generalizer.py‎

Lines changed: 22 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -6,14 +6,29 @@
66
from app.core.clauses.where_clause import CompareExpression, WhereClause
77
from app.core.schema.schema_graph import SchemaGraph
88
from app.core.schema.schema_parser import SchemaParser
9+
from app.impl.oracle_sqlpgq.schema.schema_parser import OracleSqlPgqSchemaParser
910
from app.impl.tugraph_cypher.schema.schema_parser import TuGraphSchemaParser
1011

1112

1213
class QueryGeneralizer:
13-
def __init__(self, db_id, instance_path):
14+
SCHEMA_PARSERS = {
15+
"tugraph_cypher": TuGraphSchemaParser,
16+
"tugraph": TuGraphSchemaParser,
17+
"oracle_sqlpgq": OracleSqlPgqSchemaParser,
18+
"oracle": OracleSqlPgqSchemaParser,
19+
}
20+
21+
def __init__(self, db_id, instance_path, backend: str = "tugraph_cypher"):
1422
self.db_id = db_id
1523
self.instance_path = instance_path
16-
self.schema_parser: SchemaParser = TuGraphSchemaParser(db_id, instance_path)
24+
self.backend = backend
25+
parser_class = self.SCHEMA_PARSERS.get(backend)
26+
if parser_class is None:
27+
supported = ", ".join(sorted(self.SCHEMA_PARSERS))
28+
raise ValueError(
29+
f"Unsupported schema backend '{backend}'. Supported backends: {supported}"
30+
)
31+
self.schema_parser: SchemaParser = parser_class(db_id, instance_path)
1732
self.schema_graph: SchemaGraph = self.schema_parser.get_schema_graph()
1833

1934
def generalize(self, query_pattern: List[Clause]) -> List[str]:
@@ -54,6 +69,11 @@ def generalize_from_llm(self, query_template: str) -> List[str]:
5469

5570
def generalize_from_cypher(self, query_template: str) -> List[str]:
5671
# TODO: use original awesome-text2gql to generalize new query.
72+
if self.backend not in {"tugraph_cypher", "tugraph"}:
73+
raise NotImplementedError(
74+
"generalize_from_cypher is backed by the TuGraph Cypher generalizer. "
75+
"Use get_query_pattern + generalize + an Oracle translator for oracle_sqlpgq."
76+
)
5777
from app.impl.tugraph_cypher.generalizer.graph_query_generalizer import (
5878
GraphQueryGeneralizer as CypherGeneralizer,
5979
)

‎app/core/generator/corpus_generator.py‎

Lines changed: 47 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -8,11 +8,46 @@
88

99

1010
class CorpusGenerator:
11-
def __init__(self, llm_client: LlmClient):
11+
def __init__(
12+
self,
13+
llm_client: LlmClient,
14+
query_language: str = "cypher",
15+
graph_name: str | None = None,
16+
):
1217
self.llm_client = llm_client
18+
self.query_language = query_language.lower()
19+
self.graph_name = graph_name
20+
21+
def _system_prompt(self) -> str:
22+
if self.query_language in {"oracle_sqlpgq", "sqlpgq", "sql/pgq"}:
23+
return corpus.SQLPGQ_SYSTEM_PROMPT
24+
return corpus.SYSTEM_PROMPT
25+
26+
def _instruction_template(self) -> str:
27+
if self.query_language in {"oracle_sqlpgq", "sqlpgq", "sql/pgq"}:
28+
return corpus.SQLPGQ_INSTRUCTION_TEMPLATE
29+
return corpus.INSTRUCTION_TEMPLATE
30+
31+
def _translation_prompt_template(self) -> str:
32+
if self.query_language in {"oracle_sqlpgq", "sqlpgq", "sql/pgq"}:
33+
return corpus.SQLPGQ_TRANSLATION_PROMPT_TEMPLATE
34+
return corpus.TRANSLATION_PROMPT_TEMPLATE
35+
36+
def _query_template_instruction(self) -> str:
37+
if self.query_language in {"oracle_sqlpgq", "sqlpgq", "sql/pgq"}:
38+
return corpus.SQLPGQ_QUERY_TEMPLATE_INSTRUCTION
39+
return corpus.QUERY_TEMPLATE_INSTRUCTION
40+
41+
def _query_archetypes(self) -> List[str]:
42+
if self.query_language in {"oracle_sqlpgq", "sqlpgq", "sql/pgq"}:
43+
return corpus.SQLPGQ_QUERY_ARCHETYPES
44+
return corpus.QUERY_ARCHETYPES
1345

1446
def _extract_json_from_response(self, response: str, expect_list: bool = True):
1547
"""Extract JSON from LLM response."""
48+
if not response:
49+
print(" [Warning] Empty LLM response.")
50+
return [] if expect_list else {}
1651
try:
1752
start_char, end_char = ("[", "]") if expect_list else ("{", "}")
1853
json_start = response.find(start_char)
@@ -40,7 +75,7 @@ def generate_questions_batch(
4075
all_questions = set()
4176

4277
# Randomly select a query intent archetype to guide generation
43-
archetype = random.choice(corpus.QUERY_ARCHETYPES)
78+
archetype = random.choice(self._query_archetypes())
4479
print(f"Brainstorming questions with intent: '{archetype.split(':')[0]}'")
4580

4681
instruction = corpus.EXPLORATION_PROMPT_TEMPLATE.format(
@@ -50,7 +85,7 @@ def generate_questions_batch(
5085
num_to_generate=questions_per_call,
5186
)
5287
message = [
53-
{"role": "system", "content": corpus.SYSTEM_PROMPT},
88+
{"role": "system", "content": self._system_prompt()},
5489
{"role": "user", "content": instruction},
5590
]
5691

@@ -69,16 +104,17 @@ def generate_translation_batch(
69104
self, schema_json: str, questions: List[str], error_context: Dict[str, str] = None
70105
) -> List[Dict[str, Any]]:
71106
"""
72-
Translate a list of questions into Cypher queries.
107+
Translate a list of questions into the configured graph query language.
73108
Supports retries by providing an error_context.
74109
"""
75-
instruction = corpus.TRANSLATION_PROMPT_TEMPLATE.format(
110+
instruction = self._translation_prompt_template().format(
76111
schema_json=schema_json,
77112
question=questions[0], # Assuming one question per call for clarity
113+
graph_name=self.graph_name or "GRAPH_NAME",
78114
error_context=error_context if error_context else "",
79115
)
80116
message = [
81-
{"role": "system", "content": corpus.SYSTEM_PROMPT},
117+
{"role": "system", "content": self._system_prompt()},
82118
{"role": "user", "content": instruction},
83119
]
84120

@@ -193,13 +229,14 @@ def run_generation_loop(
193229
selected_contexts = random_examples
194230

195231
# 1. Build Prompt
196-
instruction = corpus.INSTRUCTION_TEMPLATE.format(
232+
instruction = self._instruction_template().format(
197233
schema_json=schema_json,
198234
examples_json=json.dumps(selected_contexts, indent=2, ensure_ascii=False),
199235
num_per_iteration=num_per_iteration,
236+
graph_name=self.graph_name or "GRAPH_NAME",
200237
)
201238
message = [
202-
{"role": "system", "content": corpus.SYSTEM_PROMPT},
239+
{"role": "system", "content": self._system_prompt()},
203240
{"role": "user", "content": instruction},
204241
]
205242

@@ -274,7 +311,7 @@ def generate_template_based_corpus(
274311
# 3. Construct the Prompt
275312
# We directly provide the "raw" data and ask the LLM to do three things:
276313
# extract information, fill the template, and generate questions.
277-
instraction = corpus.QUERY_TEMPLATE_INSTRUCTION.format(
314+
instraction = self._query_template_instruction().format(
278315
raw_data_str=raw_data_str,
279316
current_batch_size=current_batch_size,
280317
selected_templates=selected_templates,
@@ -283,7 +320,7 @@ def generate_template_based_corpus(
283320
message = [
284321
{
285322
"role": "system",
286-
"content": "You are a helpful assistant that generates Cypher datasets.",
323+
"content": self._system_prompt(),
287324
},
288325
{"role": "user", "content": instraction},
289326
]

0 commit comments

Comments
 (0)