You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 7cd94df
Browse filesBrowse the repository at this point in the historyBrowse files
Add Oracle SQL/PGQ support to Awesome-Text2GQL (#68)
* feat: Add OracleDB support and related tests
- Updated pyproject.toml to include oracledb dependency.
- Introduced new test suite for translating Cypher queries to Oracle SQL/PGQ.
- Implemented dataset preparation tests for Oracle integration.
- Added live tests for OracleDB client functionality.
- Created query generalizer and template instantiator for Oracle SQL/PGQ.
- Enhanced corpus combiner to handle Oracle-specific queries and validation.
- Included schema parser for generating Oracle DDL statements.
* feat: Enhance Oracle SQL PGQ Translator with primary key mapping and strict validation
- Added support for node and edge primary key mappings in OracleSqlPgqQueryTranslator.
- Introduced strict property validation to ensure properties are defined for variables.
- Updated methods to normalize label maps and handle aggregate functions in WITH clauses.
- Enhanced validation for translated queries, including handling of string predicates and label predicates.
- Improved error handling for missing properties when strict validation is enabled.
- Added new command-line arguments for validation timeout and fetch limit in dataset preparation.
- Updated tests to cover new features, including primary key mapping and strict validation scenarios.
* feat: Add dataset preparation and Oracle vs Neo4j comparison utilities
* Add tests for unsupported query features and failure analysis
- Enhance `test_detect_unsupported_oracle_sqlpgq_features` with additional assertions for various unsupported query patterns.
- Introduce `test_failure_analysis_groups_unsupported_query_shapes` to analyze failure signatures for unsupported queries.
- Implement `test_failure_analysis_uses_manifest_for_invalid_schema` to validate schema direction and property checks against a manifest.
- Add normalization tests in `test_compare_normalizes_temporal_strings_and_numeric_precision` and `test_compare_normalizes_oracle_and_neo4j_node_identity`.
- Create tests for path normalization in `test_compare_normalizes_single_neo4j_path_to_flat_element_sequence`.
- Include checks for nondeterministic limits in `test_compare_detects_nondeterministic_limit_without_order_by`.
- Expose file stem label aliases in `test_loader_exposes_file_stem_label_aliases`.
* Enhance optional match handling and support for correlated optional matches
- Introduced `is_supported_correlated_optional_match` to validate correlated optional matches in Cypher queries.
- Updated `detect_unsupported_features` to remove "optional_match" feature if correlated optional matches are supported.
- Removed redundant optional match translation logic from `cypher2oracle_sqlpgq`.
- Added comprehensive tests for various optional match scenarios, including correlated optional matches and their translations to SQL.
- Improved handling of optional match clauses in the dataset preparation and query translation processes.
* feat: add CypherSchema class for schema validation and property management
- Implemented CypherSchema to manage and validate graph schema based on provided configuration.
- Added methods for detecting validation issues in Cypher queries, including node and edge label checks, property validation, and unsafe numeric conversions.
- Introduced utility functions for parsing Cypher variable labels, property references, and edge relationships.
- Included comprehensive handling of schema name aliases and property types.
- Ensured deduplication of validation issues for cleaner output.
* feat: enhance Oracle SQL PGQ Translator with stage expression correlation and numeric tolerance checks
* Enhance Cypher to Oracle SQL/PGQ translation and validation
- Introduced checks for unique schema ownership of properties in CypherSchema.
- Added detection for unsafe temporal arithmetic in aggregate queries.
- Improved handling of broad bounded variable length relationships in translation.
- Updated tests to cover new features and edge cases, including disambiguation of complex aggregate property aliases.
- Refactored unsupported feature detection to exclude expensive variable length paths.
- Enhanced query translation to preserve real ID properties over pseudo identities.
- Added stable tiebreakers for ordered queries with limits in comparison functions.
* feat: enhance detection of unsupported features with new patterns and tests
* feat: add exporter for validated Oracle SQL/PGQ dataset and enhance README
* feat: update README and dataset preparation documentation for Oracle SQL/PGQ support
Copy file name to clipboardExpand all lines: README.md
+10-1Lines changed: 10 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -37,7 +37,7 @@ Awesome-Text2GQL is an AI-assisted framework for Text2GQL dataset construction.
37
37
38
38
### Generated Benchmark Dataset
39
39
40
-
The [Text2GQL-Bench](https://arxiv.org/abs/2602.11745)'s dataset is generated by Awesome-Text2GQL framework. It contains 178,184 (Question, Query) pairs spanning 13 domains. The dataset is available at [Text2GQL-Bench_dataset](https://tugraph-web.oss-cn-beijing.aliyuncs.com/tugraph/datasets/text2gql/Text2GraphQueryBenchmark/Text2GQL-Bench_dataset.zip). To run Text2GQL test, please refer to our [Text2GraphQuery-Driver](https://github.com/TuGraph-family/text2graphquery-driver/tree/main).
40
+
The [Text2GQL-Bench](https://arxiv.org/abs/2602.11745)'s dataset is generated by Awesome-Text2GQL framework. It contains 178,184 (Question, Query) pairs spanning 13 domains. The dataset is available at [Text2GQL-Bench_dataset](https://tugraph-web.oss-cn-beijing.aliyuncs.com/tugraph/datasets/text2gql/Text2GraphQueryBenchmark/Text2GQL-Bench_dataset.zip). The dataset including the Oracle SQL/PGQ translated queries is available at [Dataset-with-SQL/PGQ](https://objectstorage.us-ashburn-1.oraclecloud.com/p/8dIkuVGsfnRQlP3ifxVDjQP0pmidpadEY18ltEbkPC4PrZyLTxjJdqDjbtWIEYUW/n/ogcs/b/Text2GQL-Bench_dataset/o/Text2GQL-Bench_dataset.zip), it includes 19633 out of 22407 existing queries. To run Text2GQL test, please refer to our [Text2GraphQuery-Driver](https://github.com/TuGraph-family/text2graphquery-driver/tree/main).
41
41
42
42
## Demo: TuGraph-DB ChatBot
43
43
@@ -195,6 +195,15 @@ After all, run:
195
195
196
196
When the script finishes, the generated corpus will be saved to examples/generated_corpus/{graph_name}_template_corpus.json.
197
197
198
+
#### Oracle SQL Property Graphs (SQL/PGQ)
199
+
200
+
Awesome-Text2GQL includes Oracle SQL/PGQ support for schema conversion, graph setup, query translation, corpus generation, validation, and benchmark dataset preparation.
201
+
202
+
For detailed workflows, see:
203
+
204
+
-[Oracle SQL/PGQ data generation workflow](./doc/en-us/development/oracle_sqlpgq_data_generation_workflow.md): convert framework/TuGraph-style schemas into Oracle SQL/PGQ artifacts, create local Oracle property graphs, generate deterministic and LLM-based corpora, validate generated queries, and combine corpus outputs.
205
+
-[Dataset preparation utilities](./dataset_prep/README.md): translate benchmark Cypher/GQL-like records to Oracle SQL/PGQ, optionally validate them against Oracle, analyze failures, compare Oracle SQL/PGQ results with Neo4j, and export validated datasets.
0 commit comments