Skip to content

Latest commit

 

History

History
168 lines (135 loc) · 4.57 KB

File metadata and controls

168 lines (135 loc) · 4.57 KB

Data Format

Normalized Graph JSON

The canonical graph format produced by the normalizer and consumed by the dataset builder:

{
  "nodes": [
    {
      "id": "com.example.MyClass",
      "type": "CLASS",
      "name": "MyClass",
      "comment": "Main service class",
      "file": "src/main/java/com/example/MyClass.java",
      "synthetic": false
    },
    {
      "id": "com.example.MyInterface",
      "type": "INTERFACE",
      "name": "MyInterface",
      "comment": "",
      "file": "src/main/java/com/example/MyInterface.java",
      "synthetic": false
    }
  ],
  "edges": [
    { "src": "com.example.MyClass", "tgt": "com.example.MyInterface", "type": "IMPLEMENTATION" },
    { "src": "com.example.MyClass", "tgt": "com.example.Utils", "type": "ASSOCIATION" }
  ],
  "metrics": {
    "com.example.MyClass": {
      "wmc": 12.0,
      "dit": 2.0,
      "noc": 3.0,
      "ac": 8.0,
      "ec": 2.0,
      "encapsulation": 0.8
    }
  }
}

Node fields

Field Type Description
id string Fully qualified name (e.g. com.example.MyClass.myMethod)
type string One of: CLASS, INTERFACE, ENUM, METHOD, FIELD, ANNOTATION, OTHER
name string Short name (e.g. myMethod)
comment string Leading doc comment, truncated to 200 chars
file string Relative source file path
synthetic bool True for auto-generated module nodes (Python/TS only)

Edge fields

Field Type Description
src string Source node ID
tgt string Target node ID
type string One of: EXTENSION, IMPLEMENTATION, COMPOSITION, AGGREGATION, ASSOCIATION

Metrics fields

Key Type Description
wmc float Weighted Methods per Class
dit float Depth of Inheritance Tree
noc float Number of Children
ac float Afferent Coupling
ec float Efferent Coupling
encapsulation float Encapsulation score (0–1)

PyG HeteroData

The normalized JSON is converted to a PyTorch Geometric HeteroData object:

data = HeteroData()

# Node features: (N, 404) float32
data["node"].x = torch.tensor(...)        # 404-dim features
data["node"].node_ids = ["id1", ...]      # original node IDs

# Edge indices per type: (2, E) long
data["node", "EXTENSION", "node"].edge_index = torch.tensor(...)
data["node", "IMPLEMENTATION", "node"].edge_index = torch.tensor(...)
data["node", "COMPOSITION", "node"].edge_index = torch.tensor(...)
data["node", "AGGREGATION", "node"].edge_index = torch.tensor(...)
data["node", "ASSOCIATION", "node"].edge_index = torch.tensor(...)

Only edge types with actual edges are present. Missing types are handled by padding with self-loops in the encoder.

Text Embeddings

Stored as compressed NumPy archives:

# data/corpus/embeddings/{repo_name}.npz
np.savez_compressed(path,
    ids=["com.example.MyClass", ...],       # (N,) string array
    embeddings=np.array([[0.1, ...], ...])  # (N, 384) float32
)

Generated by scripts/compute_embeddings.py using sentence-transformers/all-MiniLM-L6-v2 on {name} {comment}.

Training Datasets

Subgraph samples stored as pickled PyG HeteroData:

data/dataset/
├── java_spring-framework_clustered_0.pt
├── java_spring-framework_clustered_1.pt
├── java_spring-framework_random_0.pt
├── python_django_clustered_0.pt
└── ...

Each file contains one sampled subgraph (a subset of nodes and their induced edges from a single repo).

ONNX Model

Heterogeneous HGT encoder (manual_hgt.py export)

Inputs:
  node_features:    float32 (N, 404)
  edge_index_extension:      int64 (2, E1)
  edge_index_implementation: int64 (2, E2)
  edge_index_composition:    int64 (2, E3)
  edge_index_aggregation:    int64 (2, E4)
  edge_index_association:    int64 (2, E5)

Output:
  node_embeddings:  float32 (N, 128)

Homogeneous GCN scorer (to_onnx.py export)

Inputs:
  node_features:  float32 (N, 404)
  adj_matrix:     float32 (N, N)   # D^{-1/2}AD^{-1/2} normalized

Output:
  anomaly_scores: float32 (N,)     # per-node scores in [0, 1]

Model Metadata

metadata.json bundles schema information for the scoring pipeline:

{
  "schemaVersion": 1,
  "edgeTypes": ["EXTENSION", "IMPLEMENTATION", "COMPOSITION", "AGGREGATION", "ASSOCIATION"],
  "nodeFeatureDim": 404,
  "textEmbeddingDim": 384,
  "metricVectorDim": 9,
  "typeOneHotDim": 7,
  "languageOneHotDim": 3,
  "componentTypes": ["CLASS", "INTERFACE", "ENUM", "METHOD", "FIELD", "ANNOTATION", "OTHER"],
  "languages": ["java", "python", "typescript"]
}