AI & Data Topics
101 topics
- AI & Data
Windowing and Time Semantics in Stream Processing (Event Time, Watermarks, Exactly-Once)
Groups unbounded data into finite units via windowing, determines when to finalize aggregation using event time and watermarks, and guarantees exactly-once via checkpoints and transactional sinks—designing the trade-off among latency, accuracy, and completeness.
- AI & Data
Automated Machine Learning (AutoML)
A technology that automates preprocessing, model selection, HPO, neural architecture search, and ensembling as a search/optimization problem to raise productivity and accessibility, combining Bayesian/evolutionary/gradient-based search with efficient performance estimation while controlling the trade-offs of overfitting, compute cost, and explainability through governance.
- AI & Data
Central Limit Theorem, t-test, and z-test
The CLT (normal approximation of the sample mean) and the concepts, differences, and test statistics of the z-test and t-test depending on whether the population variance is known and on sample size.
- AI & Data
Graph Neural Network (GNN)
A neural network in which nodes repeatedly aggregate neighbor information (message passing) to learn representations — evolving through aggregation variants such as GCN·GraphSAGE·GAT and excelling at problems where the graph is essential, such as recommendation, drug discovery, and fraud detection.
- AI & Data
Generative AI Evaluation and Quality Management with LLM-as-a-Judge
An essay-style guide to evaluating generative AI against real business objectives and risk criteria. Covers layered model, component, and end-to-end evaluation; correctness, retrieval, groundedness, safety, and operational metrics; LLM-as-a-Judge bias and human calibration; RAG and agent testing; regression and production feedback loops; and cases in internal knowledge search, customer support, and code generation, with HELM and NIST AI RMF connections.
- AI & Data
Large Language Model Inference Optimization: KV Cache, Continuous Batching, and Speculative Decoding
An essay-style explanation of SLO-driven techniques for improving large language model serving. Covers prefill and decode; KV-cache memory estimation, paged allocation, sharing, quantization, and offloading; continuous batching, chunked prefill, and fair scheduling; and comparison of weight quantization, parallelism, and speculative decoding. Discusses customer support, internal RAG, and code-generation cases, TTFT/ITL and cost measurement, and professional-engineer considerations for cache isolation, disaggregated serving, operations, and security.
- AI & Data
Multicollinearity
A phenomenon where strong correlation among independent variables destabilizes regression coefficient estimates — diagnosis via VIF and condition index, and remedies such as variable removal, PCA, and regularization.
- AI & Data
TF-IDF (Term Importance Weighting)
The calculation process, examples, characteristics, and limitations of a technique that computes term importance as TF (frequency within a document) × IDF (log N/df).
- AI & Data
Anomaly Detection
An essay-style overview of anomaly detection, which learns the distribution and patterns of normal data to identify rare, abnormal observations outside that boundary. Covers point/contextual/collective anomaly types and the detection pipeline; the classification and comparison of statistical (Z·IQR), distance-density (LOF), isolation (Isolation Forest), and deep-learning (autoencoder·LSTM) techniques; cases in financial FDS, predictive maintenance, and AIOps/UEBA; evaluation metrics for imbalanced data (PR-AUC) and the false-positive/false-negative trade-off; and professional-engineer considerations such as explainability and concept-drift response.
- AI & Data
Equipment Predictive Maintenance Using LangChain
The concept and necessity of equipment predictive maintenance (PdM), the makeup of LangChain and LLMs (chains, agents, tools, memory, RAG), an approach that divides roles between anomaly detection (ML) and LLM interpretation to support root-cause analysis and maintenance guidance, and hallucination/security considerations.
- AI & Data
The Evolution from PLM to LLM
How PLMs, transfer-learning models based on pre-training and fine-tuning, evolve into LLMs with emergent capabilities through pre-training → SFT → RLHF/DPO alignment.
- AI & Data
Process Mining for Business Process Analysis and Improvement
An essay-style treatment of process mining as a data-driven method for discovering actual business flows from event logs, checking deviations from reference models, and improving processes by time, resource, and data perspectives—covering case/activity/event structures, discovery/conformance/enhancement, IEEE XES 1849-2023, Petri nets/BPMN/DFGs, fitness and precision, a hypothetical e-commerce approval case, and governance of log quality, privacy, and outcomes.
- AI & Data
EU AI Act: Risk Classification and Compliance Design
An essay-style treatment of EU AI Act risk classification and graduated duties for prohibited practices, transparency, high-risk systems, and general-purpose AI, with AI inventory, provider/deployer roles, lifecycle controls for data and models, human oversight, conformity evidence, incident response, recruitment and generative customer-service cases, and a method for checking the 2026 transition timeline.
- AI & Data
Image Data Annotation
Image labeling types such as bounding box, polygon, segmentation, keypoint, and 3D, with manual, semi-automatic, active-learning, and foundation-model techniques, plus IAA quality control, bias, and privacy considerations.
- AI & Data
Large-Scale Distributed Data Processing with MapReduce
An essay-style treatment of the MapReduce programming model and Hadoop execution structure for parallel processing of large key-value input through Map, Combine, Partition, Shuffle/Sort, and Reduce, including data locality, task retry, combiners, partitioners, skew, shuffle optimization, WordCount, inverted indexes, log aggregation, and comparison with Spark, distributed SQL, and stream processing.
- AI & Data
AI Model Risk Management (MRM)
An essay-style treatment of lifecycle MRM that controls losses from AI model error, misuse, change, and data bias through inventory, risk tiers, independent validation, approval, monitoring, change management, and rollback; its relationship with MLOps and AI governance; financial advisory and manufacturing inspection cases; linkage with the NIST AI RMF and ISO/IEC 42001; and an implementation roadmap from a Professional Engineer's perspective.
- AI & Data
AI TRiSM (AI Trust, Risk and Security Management) for Trustworthy AI Operations
An essay-style treatment of AI TRiSM as an integrated approach to trust, risk, and security across the AI lifecycle: its governance structure; lineage management for data, models, prompts, and agent tools; controls for fairness, explainability, prompt injection, hallucination, privacy, and drift; linkage with the NIST AI RMF and ISO/IEC 42001; financial advisory and manufacturing inspection cases; and risk-based operating strategies from a Professional Engineer's perspective.
- AI & Data
Embedding (Vector Representation)
Embedding is a representation-learning technique that maps discrete objects such as words, sentences, and images into low-dimensional dense vectors preserving semantic similarity as distance; it covers principles, types (static/contextual/multimodal), and similarity metrics, powers semantic search, RAG, and recommendation, and requires managing trade-offs of dimension, bias, drift, and privacy.
- AI & Data
Foundation Model
The concept, characteristics, and underlying technologies (Transformer, self-supervision, alignment) of large-scale pre-trained general-purpose AI models, along with legal, environmental, and social considerations.
- AI & Data
SOM (Self-Organizing Map)
Definition and characteristics of SOM, an unsupervised neural network based on competitive learning; its components (BMU, neighborhood function); and differences from supervised neural networks.
- AI & Data
AI Ethics and Governance Model
AI ethics principles such as fairness, transparency, and accountability, and a governance model spanning principles→organization→impact assessment→monitoring→regulatory compliance.
- AI & Data
Privacy-Preserving Data Collaboration Based on Data Clean Rooms
An essay-style overview of data clean rooms, which produce joint insights within approved matching, analysis and output policies without unrestricted sharing of raw data: their role, trust boundaries, controlled execution, differential privacy, cryptographic matching, output governance, applications in advertising, finance and manufacturing, interoperability, and design considerations from a Professional Engineer's perspective.
- AI & Data
AI Transparency Based on AI Model Cards and Dataset Datasheets
An overview of how to achieve AI transparency, accountability, change management, and operational control by linking model cards, which record an AI model's purpose, performance, limitations, and risks, and datasheets, which describe a dataset's provenance, composition, quality, and rights, with data lineage, the model registry, evaluation, and monitoring.
- AI & Data
Data Spaces and Data-Sovereignty-Based Trusted Data Sharing
An overview of a federated data architecture that shares data across organizations while data providers retain control, by combining trust services, catalogs, semantics, contracts, and usage policies, along with its governance, security, and diffusion strategy.
- AI & Data
Dimensionality Reduction
The concept of dimensionality reduction to mitigate the curse of dimensionality, feature selection and extraction techniques (PCA, LDA, t-SNE, autoencoders), and purpose-driven selection strategies and trade-offs.
- AI & Data
DSML Projects and MLOps
The cyclical DSML project lifecycle (business → data → modeling → operations), and MLOps that versions data, models, and code while taming model decay through pipeline automation, monitoring, and continuous training (CT).
- AI & Data
Recommendation System
An information-filtering system that predicts and ranks items a user will prefer amid information overload — combining content-based, collaborative filtering, and hybrid (deep learning) approaches in a multi-stage pipeline of candidate generation, ranking, and re-ranking.
- AI & Data
Recurrent Neural Networks (RNN) and LSTM/GRU
The structure, principles, and comparison of RNNs, which process variable-length sequences by recurrently passing past context through a hidden state, and of LSTM/GRU, which overcome vanishing gradients through gates and cell state to learn long-term dependencies, and the trade-offs versus Transformers and state-space models.
- AI & Data
Data Mining: K-means, DBSCAN, SVM
The principles, parameters, and comparison of unsupervised clustering with K-means (centroid, distance) and DBSCAN (density) and supervised classification with SVM (maximum-margin hyperplane), along with practical selection criteria.
- AI & Data
Data Visualization
A technique for conveying insight by mapping data to visual encodings with high perceptual accuracy according to the principles of accuracy, clarity, efficiency, and aesthetics, and selecting purpose-specific charts (comparison, trend, composition, relationship, distribution, spatial)—covering chartjunk removal, interactivity, accessibility, visualization ethics, and NL2Chart trends.
- AI & Data
Machine Learning Model Registry and Approval/Deployment Governance
A strategy that centrally manages machine learning models' versions, training runs, data and code lineage, evaluation, approval, and deployment pointers, and manages models as trustworthy operational assets through quality gates, canaries, rollback, audit, and supply-chain security.
- AI & Data
Synthetic Data
Artificially generated data that mimics the statistical properties of the original through rules, simulation, and generative models (GAN, VAE, Diffusion, LLM), enabling privacy protection, data augmentation, and reinforcement of sparse scenarios, while managing the fidelity, utility, and privacy trade-off through differential privacy and continuous verification.
- AI & Data
Big Data Platform Architecture Design
A scalable layered structure supporting ingestion, storage, processing, and analysis. Distributed infrastructure, data lake/warehouse/lakehouse, batch and real-time (Lambda/Kappa), and evolution toward MLOps.
- AI & Data
AI Training-Data Poisoning Attacks and Data Integrity Defense
An overview of the types and threat models of poisoning attacks that insert or tamper with malicious content in training, fine-tuning, embedding, and feedback data, and a defense strategy based on provenance, versioning, access control, behavioral verification, and rollback.
- AI & Data
AI Model Drift and Data Drift Management
An AI full-lifecycle strategy that distinguishes changes in inputs, ground truth, and predictions and performance degradation in the production environment as data, concept, and prediction drift, and manages them through baselines, statistical metrics, label delay, retraining, and rollback.
- AI & Data
Data Mining: Difference from Statistics, Structured/Unstructured, Opinion/Text
The difference between data mining and statistics, a comparison of structured and unstructured mining, and the opinion mining process versus text mining.
- AI & Data
AI Agent Orchestration and Autonomous Task Execution Governance
An orchestration structure that coordinates goal decomposition, agent selection, tool execution, state management, and verification, and the design of authorization, approval, audit, and safety governance for autonomous execution.
- AI & Data
Decision Tree
An interpretable model that builds a tree of if-then rules through repeated splitting on attribute criteria, covering information gain, Gini, overfitting, and ensembles.
- AI & Data
Retrieval-Augmented Generation (RAG)
An essay-style overview of RAG, which retrieves external documents and data at query time and injects them into LLM generation — its reference architecture, ingestion, chunking, embedding, hybrid search, re-ranking, and citation verification, a comparison of basic, advanced, and agentic types, and evaluation, freshness, access control, prompt injection countermeasures, and adoption strategy.
- AI & Data
Quality Management of AI Training Datasets
Quality assurance activities and procedures for each life-cycle stage of AI training datasets: collection→processing/labeling→inspection→utilization.
- AI & Data
Convolutional Neural Network (CNN)
A neural network that hierarchically extracts local features of an image via convolution and pooling — a standard image-recognition architecture that reduces parameters through local receptive fields and weight sharing and secures translation invariance.
- AI & Data
Data-Centric AI and Training-Data Quality Engineering
Organizes the concept, quality metrics, lifecycle, governance, and application strategy of DCAI, which systematically improves the quality, representativeness, labels, and lineage of data rather than model architecture to raise AI performance and reliability.
- AI & Data
Full-Lifecycle Information Linkage Based on the Digital Thread
An essay-style treatment of the digital thread, which links requirements, design, manufacturing, inspection, operation, and maintenance data through common identifiers, standard semantics, versions, and lineage — covering its concept, reference architecture, impact analysis, standardization, quality/security/governance, and application to manufacturing and digital twins.
- AI & Data
Prompt Engineering
A low-cost technique that controls LLM output by designing the input (role, context, examples, format) without retraining the model — key techniques such as Zero/Few-shot, CoT, and ReAct, a comparison with RAG and fine-tuning, and prompt injection security and operational governance.
- AI & Data
Metaverse Ethics Principles
The metaverse ethics principles: three core values (an intact self, safe experience, sustainable prosperity), eight principles of practice, application procedures and shared responsibility, a comparison with internet and AI ethics, and immersive-data privacy.
- AI & Data
Machine Learning Optimization Algorithms
The types, pros and cons of gradient-descent family algorithms (SGD, Momentum, AdaGrad, RMSProp, Adam) that minimize loss.
- AI & Data
Voice Data Mining
A technique that exhaustively extracts information and patterns from unstructured voice via a pipeline of STT, speaker/emotion recognition, NLP, and pattern analysis — covering processing methods, application areas, the direction of generative AI advancement, and voice-data privacy challenges.
- AI & Data
Data Product Design and Operation
The design and operational strategy of a data product that bundles data, metadata, code, infrastructure, and operational accountability into a single product delivering independent value for analytics, AI, and business decisions — covering data mesh domain ownership; discoverability, addressability, understandability, interoperability, and security; and the lifecycle from quality contracts, SLAs, and versioning to retirement, plus organizational adoption.
- AI & Data
Lambda Architecture and Kappa Architecture
Compares Lambda, which merges accuracy and real-time behavior through its three batch, speed, and serving layers, with Kappa, which unifies real-time processing and reprocessing through a single stream of an immutable log, and organizes consistency principles, reprocessing strategy, and lakehouse integration.
- AI & Data
Differences Between Machine Learning and Deep Learning
The relationship AI ⊃ Machine Learning ⊃ Deep Learning and a comparison of differences in feature extraction, data volume, resources, and interpretability.
- AI & Data
Digital Twin Architecture and Implementation Strategy
The concept and reference architecture of a digital twin that synchronizes an observable real-world target with its digital representation for observation, prediction, simulation, and optimization. Covers the physical, edge, integration, twin-core, model, and application layers, data contracts, real-time behavior, interoperability and security, manufacturing/building/supply-chain cases, and an ISO 23247-based composition strategy.
- AI & Data
AI Ethics Standards: Three Principles and Ten Requirements
Under the overarching principle of humanity in the national AI ethics standards (Dec 2020), the three basic principles (human dignity, the common good, purposiveness) and the ten core requirements. The shift from voluntary norms to the AI Basic Act (effective Jan 2026) and embedding ethics at the design stage.
- AI & Data
Federated Learning
The operating principles and algorithm (FedAvg) of federated learning, which aggregates only locally trained models without gathering the data, and its security and privacy technologies.
- AI & Data
Generative Adversarial Network (GAN)
A generative model that approximates the real distribution by pitting a generator that creates data against a discriminator that judges authenticity in a minimax game. Provides an in-depth treatment of adversarial learning principles, DCGAN/cGAN/CycleGAN/StyleGAN variants, FID/IS evaluation metrics, comparison with diffusion models, handling of mode collapse and training instability, and trust/governance issues such as deepfakes.
- AI & Data
Ensemble Learning: Bagging and Boosting
Techniques that combine multiple weak learners into a strong model. Bagging (parallel, lowers variance) and boosting (sequential, lowers bias) are complementary.
- AI & Data
LLM-as-a-Judge (LLM-Based Evaluation)
A method for using an LLM as an evaluator to measure accuracy, relevance, groundedness, and safety with rubrics — covering deterministic metrics, human meta-evaluation, bias control, RAG and agent evaluation, and operational deployment gates.
- AI & Data
Comparison of Independent-Samples and Paired-Samples t-Tests
Tests of the mean difference between two different groups (independent samples) and between paired difference values from the same subjects (paired samples). The paired-samples test cancels out individual variation for higher statistical power; selection considers sample structure, assumptions, and effect size.
- AI & Data
Bernoulli Distribution and Geometric Distribution
The definitions, expected values, variances, and memorylessness of the distributions for success/failure in a single trial (Bernoulli) and the number of trials until the first success (geometric), and their extension to the binomial and negative binomial distributions.
- AI & Data
Data Observability and Data Reliability Management
A data-reliability operations framework that collects freshness, completeness, validity, distribution, and lineage signals to detect data anomalies and connects them to impact analysis, recovery, and post-incident improvement.
- AI & Data
Overfitting
The causes of overfitting (complexity, insufficient data, noise, over-training) and remedies such as regularization, dropout, early stopping, and cross-validation, with an in-depth look at the bias-variance trade-off.
- AI & Data
LLMOps (Large Language Model Operations) and Generative-AI Service Lifecycle Management
This essay covers, in a professional-engineer answer format, the management of prompts, models, retrieval augmentation, tools, evaluation, observability, security, and cost, along with a continuous-improvement framework, for taking large-language-model-based applications from experimentation to production.
- AI & Data
Metaheuristics
General-purpose search strategies that find good approximate solutions to large-scale optimization problems in practical time. Includes genetic, simulated annealing, ant colony, and particle swarm methods.
- AI & Data
Data Lineage and Impact Analysis
This essay covers a data governance framework that traces the path by which data flows from source to transformation, storage, analysis, and service as a Dataset-Job-Run graph and analyzes change impact, root cause of failures, and personal-data propagation.
- AI & Data
ISO/IEC 42001:2023 Artificial Intelligence Management System (AIMS)
This essay covers the requirements of an AI management system (AIMS) for developing, providing, and using AI responsibly, along with risk and impact assessment, lifecycle controls, and an implementation roadmap.
- AI & Data
Principles for Selecting Big Data Analysis Tools
Principles for choosing a tool combination suited to the situation by comprehensively considering analysis purpose → data characteristics → scalability and performance → TCO, workforce competency, and governance. Also reflects the shift to cloud-managed services and the lakehouse.
- AI & Data
Feature Store and Machine Learning Data Operations
This essay covers the components of a feature store that centralizes the definition, transformation, validation, and lineage of machine-learning features and serves them consistently to offline training and online inference, along with point-in-time correctness and data-leakage prevention, batch and streaming operations, and quality, security, and governance strategies, in a professional-engineer essay format.
- AI & Data
GraphRAG (Graph Retrieval-Augmented Generation)
The structure, retrieval methods, construction process, and evaluation and operational strategy of GraphRAG, which extracts entities, relationships, and communities from unstructured documents and combines hierarchical summarization with graph traversal to overcome the local-retrieval limits of ordinary RAG.
- AI & Data
Characteristics of the Normal Distribution
A continuous, bell-shaped distribution symmetric about its mean. As a foundation of statistical inference through the 68-95-99.7 rule, standardization (Z), and the central limit theorem, it also addresses fat-tail risk and normality testing.
- AI & Data
On-Device AI
The concept of on-device AI, which performs AI inference directly on the device, its HW/SW technologies (NPU, model compression), and a comparison with cloud AI.
- AI & Data
MCP (Model Context Protocol)
The structure, primitives, transport, and security threats of MCP — an open protocol that connects LLM applications with external tools, data, and prompts in a standardized way through a Host-Client-Server structure and JSON-RPC 2.0, reducing the M×N integration problem to M+N — along with professional-engineer implications.
- AI & Data
Correlation and Causation
A comparison of correlation, where things vary together, and causation, a cause-and-effect relationship; the pitfall that correlation ≠ causation; and verification through experiments and causal inference.
- AI & Data
Data Contracts and Data Product Reliability
A data-product interface in which data producers and consumers explicitly agree on the meaning, schema, quality, freshness, access, and service level of data and automatically validate them. Going beyond the limits of fixing only the schema, this essay covers the components of a contract, its lifecycle, compatibility and versioning strategies, dbt model contracts and ODCS integration, and operational considerations in data-mesh and DataOps environments.
- AI & Data
Digital Twin and Metaverse
A comparison of the digital twin, which replicates and predicts reality, with the metaverse for virtual experience and communication, and their convergence in the industrial metaverse.
- AI & Data
Inductive Reasoning and Machine Learning
The correspondence between inductive reasoning, which generalizes rules from individual cases, and machine learning that automates it; hypothesis space and inductive bias; the limits of generalization error and data bias; and complements such as neuro-symbolic AI, RAG, and XAI.
- AI & Data
Spiking Neural Network (SNN)
A third-generation, event-driven neural network that mimics neuronal spike firing and firing timing (temporal information). It computes only when a spike occurs, achieving ultra-low power and low latency; because backpropagation is hard to apply, it advances alongside dedicated training methods such as surrogate gradients, STDP, and ANN conversion, and neuromorphic hardware.
- AI & Data
Data Commerce
Data-driven commerce that uses AI to analyze customer data and recommends and sells personalized products and services at the right time.
- AI & Data
DeepView
DeepView, a deep-learning-based intelligent video content analysis (VCA) technology — its concept and processing pipeline (object detection, tracking/re-identification, behavior analysis, intelligent search), use cases, and advanced challenges such as edge AI, privacy, and bias.
- AI & Data
Machine Learning Modeling and ModelOps
The difference between modeling, which develops models, and ModelOps, which continuously manages models as trusted assets through deployment, monitoring, retraining (CT), and governance — their architecture, maturity, and extension to LLMOps.
- AI & Data
Diffusion Model
A generative deep-learning model composed of a forward diffusion process that progressively adds noise to data and a reverse diffusion process that undoes it with a neural network (U-Net/DiT). An essay-style overview of the principle of decomposing generation into a multi-step denoising problem to secure both training stability and quality/diversity, DDPM's forward and reverse processes and noise/loss structure, the compute-saving principles of conditional generation (Classifier-Free Guidance) and latent diffusion (LDM/Stable Diffusion), a quality/diversity/speed trade-off comparison with GANs and VAEs, evaluation metrics such as FID, and governance considerations such as copyright, deepfakes, and watermarking.
- AI & Data
Knowledge Distillation
A leading model-compression technique that transfers the soft targets (tacit knowledge, inter-class similarity) of a large, high-performing teacher model to a small student model via temperature scaling and KL-divergence loss, reducing parameters, computation, and latency while preserving accuracy as much as possible. An essay-style overview of the soft-target principle, the combination of temperature and loss, the response/feature/relation/self-distillation types, comparison and combination with pruning and quantization, and Professional-Engineer-level considerations such as the teacher-student capacity gap, data licensing, and bias transfer.
- AI & Data
Explainable AI (XAI)
A technology that presents the reasoning, contributing factors, and decision process of black-box models such as deep learning in human-understandable form (LIME, SHAP, Grad-CAM) to ensure transparency, trustworthiness, and accountability—balancing the accuracy–interpretability–cost trade-off on a purpose- and risk-based basis and forming a core pillar of AI trustworthiness and regulatory compliance such as the EU AI Act.
- AI & Data
Multimodal AI
A technology that aligns heterogeneous modalities—text, image, audio, and video—into a common semantic space (CLIP contrastive learning) and fuses them via early, late, and cross-attention to achieve complementary understanding and generation—where overcoming inter-modal heterogeneity and controlling hallucination and bias are the challenges; it is the standard form of foundation models and a key path toward VLA and embodied intelligence.
- AI & Data
Mixture of Experts (MoE)
A large-scale AI scaling architecture in which a gating network selectively activates only a few experts per token—conditional computation that decouples total parameters (knowledge capacity) from active parameters (compute), realizing a far larger model within the same compute budget—covering routing, load balancing, expert parallelism, and comparison with dense models.
- AI & Data
AI Agents and Agentic AI
A system that combines an LLM 'brain' with planning, memory, tools, and actions to autonomously repeat observe-think-act loops until a goal is achieved—spreading around ReAct, multi-agent designs, and the MCP standard, where balancing autonomy against control and safety is key.
- AI & Data
Knowledge Graph
A knowledge base that represents entities as nodes and relationships as edges in triples, treating meaning and connections as first-class citizens—built around RDF and property-graph models, ontology reasoning, and graph embeddings for search, recommendation, and fraud detection, and re-emerging as an explainable knowledge source that curbs LLM hallucination via GraphRAG.
- AI & Data
Reinforcement Learning
A paradigm in which an agent learns an optimal policy that maximizes cumulative delayed, evaluative rewards through trial-and-error interaction with an environment—covering value-based, policy-based, and actor-critic families, the exploration-exploitation balance, and extension to LLM alignment via RLHF.
- AI & Data
Transformers and the Attention Mechanism
An architecture that, without recurrence or convolution, computes relationships across an entire sequence in parallel using only Query-Key-Value self-attention, achieving long-range dependency modeling and scalability—covering multi-head attention, positional encoding, encoder/decoder structure, and O(n²) efficiency improvements.
- AI & Data
Data Mesh and Distributed Data Architecture
A socio-technical distributed data architecture that decentralizes ownership of analytical data from a central data team to business domains, treats data as a discoverable, trustworthy product (Data as a Product), and maintains interoperability through a self-service platform and policy-as-code federated governance—delving into its four principles, its differences from the lakehouse and data fabric, and incremental adoption strategies.
- AI & Data
Data Lakehouse Architecture and Governance
Organizes an architecture that combines the scalability and openness of a data lake with the consistency and manageability of a data warehouse, integrating Bronze, Silver, and Gold data products with ACID, schema, quality, lineage, and privacy governance.
- AI & Data
AI Trustworthiness
The trustworthy quality of AI, secured through whole-lifecycle governance, with fairness, transparency, robustness, safety, accountability, and privacy as its core attributes.
- AI & Data
Bayesian Optimization
A technique that balances exploration and exploitation using a surrogate model (GP) and acquisition functions (EI, UCB) to find the optimum of a black-box function in few trials.
- AI & Data
Data Exchange
A brokerage platform for registering, discovering, trading, and settling data as a commodity — catalogs, valuation, and security controls alongside privacy and quality issues.
- AI & Data
Deepfake
Technology that synthesizes realistic fake media using generative AI such as GANs, and countermeasures including detection AI, watermarking, provenance authentication, and legislation.
- AI & Data
Artificial Neural Network
A layered structure of neurons combining weighted sums and activation — FFNN forward prediction, backpropagation learning, and activation functions such as Sigmoid, ReLU, and Softmax.
- AI & Data
Point Estimation vs. Interval Estimation
Concepts and comparison of point estimation, which estimates a parameter as a single value (unbiased, efficient), and interval estimation, which estimates it as a range with a confidence level.
- AI & Data
Margin Classification in Linear SVM (Hard/Soft Margin)
Margin maximization in a linear SVM and its two classification approaches: the hard margin (no misclassification allowed) and the soft margin (misclassification allowed via slack ξ, controlled by C).
- AI & Data
AI Fine-tuning
A transfer-learning technique that further trains a pre-trained model on task-specific data — Full/PEFT (LoRA, QLoRA), SFT and RLHF/DPO alignment, and a comparison with RAG.
- AI & Data
Legal, Ethical, and Technical Issues of AI Systems and Solutions
The legal (liability, copyright), ethical (bias, transparency), and technical (hallucination, security) issues of AI systems, and solutions through governance, XAI, and MLOps.
- AI & Data
Machine Learning Performance Metrics
Definitions, formulas, and selection criteria for classification (Accuracy, Precision, Recall, F1, ROC-AUC) and regression (MAE, RMSE, R²) metrics, plus the confusion matrix and the precision-recall trade-off.
- AI & Data
Multi-GPU Technology (Large-Scale Neural Network Training)
The concept and benefits of multi-GPU, data/model/pipeline/tensor parallelism, and communication, memory, and synchronization considerations when building the environment.
- AI & Data
RAG (Retrieval-Augmented Generation)
A technique and architecture that combines external knowledge retrieval results into the LLM prompt to reduce hallucination and generate up-to-date, evidence-based answers.