Narzędzia użytkownika

Narzędzia witryny


opus31

To jest stara wersja strony!


Explainable Behavioral Engineering (NCN Opus 31/ST6)

Research Vision

This research programme investigates how behavioral models, formal logic, automated reasoning, and explainable diagnostics can support the analysis, verification, and understanding of software systems developed with AI assistance.

The long-term objective is to establish foundations for Explainable Behavioral Engineering (EBE), enabling the transformation of requirements and behavioral artifacts into formally analyzable representations supporting verification, diagnosis, and explainability.

Requirements
    ↓
Behavioral Models
    ↓
Logical Specifications
    ↓
Verification
    ↓
Diagnostic Explainability
    ↓
Behavioral Accountability

Level 1: Published Foundations

This level contains peer-reviewed results forming the published scientific foundation of the programme. ([Kli 25 FSE-AI-IDE], [Kli Sem 25 EASE], [Kli Wit 25 EMSE], [Kli 25 PACIS], [Kli 24 IS], [Kli Wit 24 ASE-RENE], [Kli 24 ASE-ASYDE], [Kli 24 ISD], [Kli 23 ISD], [Kli 19 LAMP], [Kli 18 Access], [Kli 14 AMCS].) The complete publication record underpinning these foundations is available in the author's publication profile and in the project proposal and is therefore not reproduced on this page. The published foundations establish the feasibility of:

  • transforming behavioral models into logical specifications,
  • applying automated reasoning and theorem proving,
  • integrating behavioral modeling with software engineering workflows,
  • supporting AI-assisted software engineering through formal methods.

Level 2: Ongoing Research (=articles under review)

This level documents current research directions extending the published foundations toward a coherent programme of explainable and verifiable behavioral engineering.

Behavioral Modeling and Elicitation

[Kli 26d] Radoslaw Klimek „AI-Assisted Behavioural Modelling from Logical Squares”. Title changed to „Reasoning Scaffolds for AI-Assisted Behavioural Modelling: A Logical-Square Instantiation”. Available: https://drive.google.com/file/d/13UpfESRf0Lc_BKVmmwIeuiqYNuT2GVdk/view?usp=drive_link. ACCEPTED Conference paper, MODELS 2026 (Rank A, 140 MNiSW points). Acceptance notification: https://drive.google.com/file/d/1ZCM-jHR01vMIH-sOkXksY8ikYh4AcmxU/view?usp=drive_link.

Behavioral Verification

[Kli Wit 25 EMSE] Radosław Klimek, Julia Witek “Logic Mining from Process Logs: Towards Automated Specification and Verification”. Pre-print version available at https://arxiv.org/abs/2506.08628 The extended version of the conference paper has been submitted to the JCR journal with Q2 quartiles, 140 MNiSW points, under review.

[Kli et at 26] Natania Dyczek, Jagoda Flejmer, Radoslaw Klimek „Why Formal Constraints Fail on Real-World Execution Logs: The Dead-End Phenomenon”. Title changed to „Dead-End Ratio: A Behavioral Model Quality Metric for Constraint-Aware Runtime Verification”. Available at: https://drive.google.com/file/d/1iGBrCVeVksMM0Q756ZMEzfaKy9vDM6NU/view?usp=drive_link. Workshop Flagship conference paper, Core Rank A*, 200 pkt MNiSW, under review.

Diagnostic Explainability

[Kli Bla 26] Radosław Klimek, Jakub Blazowski “Towards Diagnostic Explainability inWorkflow Verification via Shapley Attribution”. Available at: https://drive.google.com/file/d/1AtZUpwG9lA3X074wgk1LmCLyrRPdi-ih/view?usp=drive_link. ACCEPTED Conference paper, ASE 2026 (Flagship conference Core Rank A*, 200 MNiSW points). Acceptance notification: https://drive.google.com/file/d/1jR-GT5LYMa7f6Vnon_D183MkQAcbwfih/view?usp=drive_link. See conference program: https://conf.researchr.org/track/ase-2026/ase-2026-nier

[Kli 26] Radoslaw Klimek „Toward Defensible System Behavior in AI-Assisted Software Engineering”. Title changed to „Toward Defensible System Behavior for AI-Generated Software Artifacts” Available at: https://drive.google.com/file/d/1Siu6l0XGr0qsdpUewYlFhHhhwRjn_qIo/view?usp=drive_link. Workshop Flagship conference paper, Rank A*, 200 pkt pkt MNiSW, under review.

[Klo Kli 26] Michał Klos, Radoslaw Klimek „Learning System Behavior from Logs: Graph-Based Anomaly Detection Using Attribute-Aware Autoencoders”. Available at: https://drive.google.com/file/d/1FEitqDCgZxdCjjI4pXMdOEoLzQSU2cMQ/view?usp=sharing. Conference paper, Rank A, 140 pkt pkt MNiSW, under review.

AI-Assisted Software Engineering

[Kli 26b] Radoslaw Klimek „What Fails in Empirical SE and AI4SE? A Study of Evaluation Fragility”. Available at: https://drive.google.com/file/d/128JSq_4HOYNfb4QZc_73Q9TJfbSf8vfk/view?usp=drive_link (and also https://drive.google.com/file/d/1UBC-9oOMX1ga7pfL6l9g3mObFpqD1ezy/view?usp=drive_link). ACCEPTED Conference paper, ICSEM 2026 (Conference Core Rank A, 140 MNiSW points). Acceptance notification: https://drive.google.com/file/d/11Nxa5CnOSNd5-mEoSNcwEVLHqBwDGWhd/view?usp=drive_link. See conference program: https://conf.researchr.org/track/icsme-2026/icsme-2026-replication-and-negative-results?#event-overview.

[Kli 26c] Radoslaw Klimek „A Workflow-Based LLM Assistant for Iterative and Verified Requirements Engineering”. Title changed to: „A Workflow-Driven Multi-Agent Architecture for Requirements Engineering”. Available at: https://drive.google.com/file/d/16SyfChWHyUOVj-OqM9Heznidq3ytD_7R/view?usp=drive_link. Workshop Flagship conference paper, Rank A*, 200 pkt pkt MNiSW, under review.

Others / Adjacent Research

[Kli 26e] Radoslaw Klimek „Context-Aware Orchestration of Adaptive Decisions in Intelligent Environments”. Available at: https://drive.google.com/file/d/1RHOAYRuJ_aKXy0ZWlQIhk41u0z7dmdqh/view?usp=drive_link. Journal JCR/IF paper, Q1 quartile, 200 pkt, under review.

[Kli Ole 26] Radoslaw Klimek, Arkadiusz Olesek „A prole-based framework for the generation of synthetic tourist mobility trajectories in urban decision support”. Available at: https://drive.google.com/file/d/1l38kJN7Pv5zo2MUTrMkW2A6Ca5P7wuWU/view?usp=drive_link. Journal JCR/IF paper, Q1 quartile, 200 pkt, under review.

Level 3: Experimental Platforms and Research Infrastructure (future publications and ongoing developments)

This level contains experimental environments, benchmarks, datasets, and prototypes used to validate research hypotheses and demonstrate technical feasibility. This level comprises experimental environments, benchmark collections, datasets, and prototype systems developed to validate research hypotheses and demonstrate the technical feasibility of the proposed methods.

LOGIMINE

LOGIMINE ('Log'ic 'Mine'ing and Verification Environment) is a research platform implementing the concept of 'Logic Mining', an approach that combines process mining, formal methods, automated theorem proving, and AI-assisted software engineering within a unified behavioral reasoning environment. The central idea is to transform process execution data into formal logical representations that can be automatically analyzed, verified, compared, and explained using mature reasoning technologies.

Unlike traditional process mining systems, which typically stop after discovering a behavioral model, LOGIMINE continues the analysis by automatically generating formal logical specifications and submitting them to automated theorem provers. Consequently, behavioral models become executable logical objects that support rigorous reasoning about process properties, correctness, consistency, robustness, and semantic relationships between independently discovered models.

The platform provides a complete transformation pipeline starting from standard event logs (XES or CSV). Event logs are analyzed using state-of-the-art process mining techniques to discover process trees representing the observed behavior. These trees are automatically translated into workflow expressions and subsequently transformed into first-order logical specifications using the compositional Pattern-Composition (PC) algorithm. The resulting specifications are converted into the TPTP standard and verified using automated theorem provers, including Vampire and E-Prover.

Main Capabilities

LOGIMINE integrates a broad spectrum of research and engineering functionalities within a single environment.

  • 'Process Discovery' – automatic discovery of behavioral models from execution traces using process mining algorithms.
  • 'Behavioral Modeling' – representation of discovered behavior using process trees and workflow expressions suitable for further formal analysis.
  • 'Logic Mining' – automatic generation of first-order logical specifications from behavioral models through compositional logical specification generation.
  • 'Automated Theorem Proving' – formal verification of generated specifications using state-of-the-art theorem provers supporting the TPTP standard.
  • 'Behavioral Verification' – analysis of consistency, satisfiability, behavioral correctness, safety, and other logical properties of process models.
  • 'Process Comparison' – formal comparison of independently discovered models using logical equivalence, implication, inclusion, conjunction, and other semantic relations.
  • 'Compliance and What-if Analysis' – evaluation of behavioral constraints, hypothetical modifications, and alternative execution scenarios.
  • 'Visualization' – graphical presentation of process trees, workflow structures, logical specifications, and verification outcomes.
  • 'Benchmark Management' – execution and comparison of verification experiments on benchmark collections, including performance evaluation of multiple theorem provers.
  • 'Research Data Management' – organization of generated specifications, benchmark datasets, verification artifacts, and experimental results.

AI-Assisted Behavioral Analysis

LOGIMINE extends classical process mining by incorporating Large Language Models (LLMs) into the behavioral analysis workflow. Instead of replacing formal reasoning, LLM agents complement symbolic methods by generating synthetic event logs from natural-language descriptions, constructing experimental scenarios, supporting benchmark generation, and assisting users during exploratory behavioral analysis.

The platform currently supports multiple LLM providers through a unified interface, enabling comparative evaluation of different models while maintaining the same underlying formal verification pipeline. Regardless of the selected LLM, all generated behavioral artifacts are ultimately verified using symbolic reasoning rather than accepted solely on the basis of probabilistic model outputs.

Digital Twin Game

One of the most distinctive components of LOGIMINE is the 'Digital Twin Game', an adversarial verification framework combining digital twins, automated theorem proving, and LLM agents. In this setting, an LLM-based adversary actively searches for behavioral weaknesses by generating controlled attacks against a process model discovered from execution logs.

Each generated attack is translated into a formal logical problem and evaluated using automated theorem provers. Successful attacks identify behavioral vulnerabilities, whereas unsuccessful attacks demonstrate behavioral robustness. The framework additionally supports iterative self-correction by proposing model modifications and immediately re-verifying their correctness, creating a closed loop of attack, verification, repair, and validation.

System Architecture

LOGIMINE has been designed as a modular research platform consisting of interoperable components:

  • Core Logic Mining engine;
  • Process Mining and Behavioral Modeling module;
  • Logic Specification Generator;
  • Automated Theorem Proving layer;
  • Behavioral Verification framework;
  • Process Comparison module;
  • Interactive Web IDE;
  • Synthetic Event Log Generator;
  • Digital Twin Game;
  • Benchmark Manager;
  • Visualization and Analytics modules.

The modular architecture enables independent development of individual components while preserving a common behavioral reasoning pipeline shared by all analysis modes.

Research Applications

LOGIMINE serves as an experimental infrastructure for research in Logic Mining, behavioral modeling, process mining, formal verification, automated reasoning, explainable behavioral analysis, and AI-assisted software engineering. The platform supports reproducible experimental studies, comparative evaluation of verification methods, benchmark development, and investigation of novel concepts combining symbolic reasoning with generative artificial intelligence.

By integrating behavioral modeling, logical specification generation, automated theorem proving, explainable verification, and AI-assisted analysis within a single environment, LOGIMINE provides a comprehensive platform for developing trustworthy, explainable, and formally verifiable behavioral engineering methods.

A central component is LOGIMINE (Logic Mining and Verification Environment), a research platform supporting the complete pipeline from event logs to formal behavioral analysis. The platform integrates workflow discovery from execution traces, behavioral modeling using workflows and process trees, automatic generation of logical specifications, formal verification through automated theorem proving, process comparison, compliance checking, what-if analysis, and visual analytics for exploring behavioral structures and verification outcomes.

Further technical details are available in the accompanying documentation: LOGIMINE IDE Documentation, LOGIMINE System Documentation2, and Digital Twin Game Documentation.

LOFT

Another key component is LOFT (Logical Framework and Testbench), a benchmark generation and experimentation environment for automated reasoning in software engineering. LOFT supports the generation and management of logical verification problems, including satisfiability, consistency, implication, redundancy, conflict detection, behavioral constraints, and theorem-proving tasks. The framework enables the construction of benchmark families with controllable structural properties, facilitating systematic evaluation of theorem provers, SAT/SMT solvers, and logic-based verification methods.

ATP Benchmark Collection

The infrastructure further includes the ATP Benchmark Collection, a benchmark suite of automated theorem proving problems derived from software-engineering-oriented behavioral verification tasks, and the Logical Problem Catalog, a curated repository of logical verification problems involving behavioral models, workflow specifications, consistency checking, satisfiability analysis, property validation, implication reasoning, and behavioral diagnostics. Together, these platforms and resources provide an experimental foundation for advancing explainable and verifiable behavioral engineering through reproducible evaluation, benchmark-driven research, and prototype validation.

RE-IDE

Workflow-driven requirements engineering environment supporting structured model generation, clarification, validation, and preparation of artifacts for formal verification.

Programme Summary

The programme integrates software engineering, requirements engineering, formal methods, process mining, automated reasoning, and AI-assisted development. Its central objective is to move from artifact-level correctness toward explainable and verifiable reasoning about system behavior.

opus31.1785611061.txt.gz · ostatnio zmienione: przez admin