The canonical graph format produced by the normalizer and consumed by the dataset builder:
{
"nodes": [
{
"id": "com.example.MyClass",
"type": "CLASS",
"name": "MyClass",
"comment": "Main service class",
"file": "src/main/java/com/example/MyClass.java",
"synthetic": false
},
{
"id": "com.example.MyInterface",
"type": "INTERFACE",
"name": "MyInterface",
"comment": "",
"file": "src/main/java/com/example/MyInterface.java",
"synthetic": false
}
],
"edges": [
{ "src": "com.example.MyClass", "tgt": "com.example.MyInterface", "type": "IMPLEMENTATION" },
{ "src": "com.example.MyClass", "tgt": "com.example.Utils", "type": "ASSOCIATION" }
],
"metrics": {
"com.example.MyClass": {
"wmc": 12.0,
"dit": 2.0,
"noc": 3.0,
"ac": 8.0,
"ec": 2.0,
"encapsulation": 0.8
}
}
}| Field | Type | Description |
|---|---|---|
id |
string | Fully qualified name (e.g. com.example.MyClass.myMethod) |
type |
string | One of: CLASS, INTERFACE, ENUM, METHOD, FIELD, ANNOTATION, OTHER |
name |
string | Short name (e.g. myMethod) |
comment |
string | Leading doc comment, truncated to 200 chars |
file |
string | Relative source file path |
synthetic |
bool | True for auto-generated module nodes (Python/TS only) |
| Field | Type | Description |
|---|---|---|
src |
string | Source node ID |
tgt |
string | Target node ID |
type |
string | One of: EXTENSION, IMPLEMENTATION, COMPOSITION, AGGREGATION, ASSOCIATION |
| Key | Type | Description |
|---|---|---|
wmc |
float | Weighted Methods per Class |
dit |
float | Depth of Inheritance Tree |
noc |
float | Number of Children |
ac |
float | Afferent Coupling |
ec |
float | Efferent Coupling |
encapsulation |
float | Encapsulation score (0–1) |
The normalized JSON is converted to a PyTorch Geometric HeteroData object:
data = HeteroData()
# Node features: (N, 404) float32
data["node"].x = torch.tensor(...) # 404-dim features
data["node"].node_ids = ["id1", ...] # original node IDs
# Edge indices per type: (2, E) long
data["node", "EXTENSION", "node"].edge_index = torch.tensor(...)
data["node", "IMPLEMENTATION", "node"].edge_index = torch.tensor(...)
data["node", "COMPOSITION", "node"].edge_index = torch.tensor(...)
data["node", "AGGREGATION", "node"].edge_index = torch.tensor(...)
data["node", "ASSOCIATION", "node"].edge_index = torch.tensor(...)Only edge types with actual edges are present. Missing types are handled by padding with self-loops in the encoder.
Stored as compressed NumPy archives:
# data/corpus/embeddings/{repo_name}.npz
np.savez_compressed(path,
ids=["com.example.MyClass", ...], # (N,) string array
embeddings=np.array([[0.1, ...], ...]) # (N, 384) float32
)Generated by scripts/compute_embeddings.py using sentence-transformers/all-MiniLM-L6-v2 on {name} {comment}.
Subgraph samples stored as pickled PyG HeteroData:
data/dataset/
├── java_spring-framework_clustered_0.pt
├── java_spring-framework_clustered_1.pt
├── java_spring-framework_random_0.pt
├── python_django_clustered_0.pt
└── ...
Each file contains one sampled subgraph (a subset of nodes and their induced edges from a single repo).
Inputs:
node_features: float32 (N, 404)
edge_index_extension: int64 (2, E1)
edge_index_implementation: int64 (2, E2)
edge_index_composition: int64 (2, E3)
edge_index_aggregation: int64 (2, E4)
edge_index_association: int64 (2, E5)
Output:
node_embeddings: float32 (N, 128)
Inputs:
node_features: float32 (N, 404)
adj_matrix: float32 (N, N) # D^{-1/2}AD^{-1/2} normalized
Output:
anomaly_scores: float32 (N,) # per-node scores in [0, 1]
metadata.json bundles schema information for the scoring pipeline:
{
"schemaVersion": 1,
"edgeTypes": ["EXTENSION", "IMPLEMENTATION", "COMPOSITION", "AGGREGATION", "ASSOCIATION"],
"nodeFeatureDim": 404,
"textEmbeddingDim": 384,
"metricVectorDim": 9,
"typeOneHotDim": 7,
"languageOneHotDim": 3,
"componentTypes": ["CLASS", "INTERFACE", "ENUM", "METHOD", "FIELD", "ANNOTATION", "OTHER"],
"languages": ["java", "python", "typescript"]
}