Synthetic data generation

Generating synthetic data

Synthetic data generation methods can be broadly categorised into two approaches:

  1. Statistical/Tabular Methods: Using statistical distributions and predefined structures (e.g., scikit-learn methods)
  2. Generative AI Methods: Using deep learning and LLMs to learn and replicate data distributions (e.g., SDV, CTGAN, LLMs)

This page covers tabular synthetic data generation using scikit-learn. For advanced generative AI methods such as using the SDV library please check the SDV page, which includes methods such as Gaussian copulas, CTGAN and CopulaGAN.

Tabular data generation

Regression data

What does a regression consist of?

For this section we will mainly use scikit-learn’s make_regression method.

For reproducibility, we will set a random_state.

We will create a dataset using make_regression’s random linear regression model with input features x=(f1,f2,f3,f4)x=(f_1,f_2,f_3,f_4) and an output yy.

import numpy as np
import pandas as pd
from sklearn.datasets import make_regression
from scipy.stats import linregress

N_FEATURES = 4
N_TARGETS = 1
N_SAMPLES = 100
random_state = 42

dataset = make_regression(
    n_samples=N_SAMPLES,
    n_features=N_FEATURES,
    n_informative=2,
    n_targets=N_TARGETS,
    bias=0.0,
    effective_rank=None,
    tail_strength=0.5,
    noise=0.0,
    shuffle=True,
    coef=False,
    random_state=random_state,
)

print(dataset[0][:10])
print(dataset[1][:10])

Let’s turn this dataset into a Pandas DataFrame:

df = pd.DataFrame(data=dataset[0], columns=[f"f{i+1}" for i in range(N_FEATURES)])

df["y"] = dataset[1]

df.head()

Let’s plot the data:

figure

Changing the Gaussian noise level

The noise parameter in make_regression allows to adjust the scale of the data’s gaussian centered noise.

dataset = make_regression(
    n_samples=N_SAMPLES,
    n_features=N_FEATURES,
    n_informative=2,
    n_targets=N_TARGETS,
    bias=0.0,
    effective_rank=None,
    tail_strength=0.5,
    noise=2.0,
    shuffle=True,
    coef=False,
    random_state=random_state,
)

df = pd.DataFrame(data=dataset[0], columns=[f"f{i+1}" for i in range(N_FEATURES)])

df["y"] = dataset[1]
figure

Visualising increasing noise

Let’s increase the noise by 10i10^i, for i=1,2,3i=1, 2, 3 and see what the data looks like.

df = pd.DataFrame(data=np.zeros((N_SAMPLES, 1)))

def create_noisy_data(noise):
    return make_regression(
        n_samples=N_SAMPLES,
        n_features=1,
        n_informative=1,
        n_targets=1,
        bias=0.0,
        effective_rank=None,
        tail_strength=0.5,
        noise=noise,
        shuffle=True,
        coef=False,
        random_state=random_state,
    )


for i in range(3):
    data = create_noisy_data(10 ** i)

    df[f"f{i+1}"] = data[0]
    df[f"y{i+1}"] = data[1]
figure

Classification data

To generate data for classification we will use the make_classification method.

from sklearn.datasets import make_classification

N = 4
data = make_classification(
    n_samples=N_SAMPLES,
    n_features=N,
    n_informative=4,
    n_redundant=0,
    n_repeated=0,
    n_classes=2,
    n_clusters_per_class=1,
    weights=None,
    flip_y=0.01,
    class_sep=1.0,
    hypercube=True,
    shift=0.0,
    scale=1.0,
    shuffle=True,
    random_state=random_state,
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N)])
df["y"] = data[1]
df.head()
figure

Cluster separation

According to the docs1 https://scikit-learn.org/stable/modules/generated/sklearn.datasets.make_classification.html, class_sep is the factor multiplying the hypercube size.

Larger values spread out the clusters/classes and make the classification task easier.

N_FEATURES = 4

data = make_classification(
    n_samples=N_SAMPLES,
    n_features=N_FEATURES,
    n_informative=4,
    n_redundant=0,
    n_repeated=0,
    n_classes=2,
    n_clusters_per_class=1,
    weights=None,
    flip_y=0.01,
    class_sep=3.0,
    hypercube=True,
    shift=0.0,
    scale=1.0,
    shuffle=True,
    random_state=None,
)

df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N_FEATURES)])

df["y"] = data[1]
figure

We can make the cluster separability more difficult, by decreasing the value of class_sep.

N_FEATURES = 4

data = make_classification(
    n_samples=N_SAMPLES,
    n_features=N_FEATURES,
    n_informative=4,
    n_redundant=0,
    n_repeated=0,
    n_classes=2,
    n_clusters_per_class=1,
    weights=None,
    flip_y=0.01,
    class_sep=0.5,
    hypercube=True,
    shift=0.0,
    scale=1.0,
    shuffle=True,
    random_state=None,
)

df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N_FEATURES)])

df["y"] = data[1]
figure

Noise level

According to the documentation2 https://scikit-learn.org/stable/modules/generated/sklearn.datasets.make_classification.html, flip_y is the fraction of samples whose class is assigned randomly.

Larger values introduce noise in the labels and make the classification task harder.

figure
df = pd.DataFrame(data=np.zeros((N_SAMPLES, 1)))

for i in range(3):
    data = make_classification(
        n_samples=N_SAMPLES,
        n_features=2,
        n_informative=2,
        n_redundant=0,
        n_repeated=0,
        n_classes=2,
        n_clusters_per_class=1,
        weights=None,
        flip_y=0,
        class_sep=i + 0.5,
        hypercube=True,
        shift=0.0,
        scale=1.0,
        shuffle=False,
        random_state=random_state,
    )
    df[f"f{i+1}1"] = data[0][:, 0]
    df[f"f{i+1}2"] = data[0][:, 1]
    df[f"t{i+1}"] = data[1]
figure

It is noteworthy that many paremeters in scikit-learn for synthetic data generation allow inputs per feature or cluster. To do so, we simple pass the parameter value as an array. For instance, to

N = 4

data = make_classification(
    n_samples=N_SAMPLES,
    n_features=N,
    n_informative=4,
    n_redundant=0,
    n_repeated=0,
    n_classes=2,
    n_clusters_per_class=1,
    weights=None,
    flip_y=0.01,
    class_sep=1.0,
    hypercube=True,
    shift=0.0,
    scale=1.0,
    shuffle=True,
    random_state=random_state,
)

df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N)])

df["y"] = data[1]

Cluster separability

from sklearn.datasets import make_blobs

N_FEATURE = 4
data = make_blobs(
    n_samples=60,
    n_features=N_FEATURE,
    centers=3,
    cluster_std=1.0,
    center_box=(-5.0, 5.0),
    shuffle=True,
    random_state=None,
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N_FEATURE)])
df["y"] = data[1]
figure

To make a cluster more separable we can change cluster_std.

data = make_blobs(
    n_samples=60,
    n_features=N_FEATURES,
    centers=3,
    cluster_std=0.3,
    center_box=(-5.0, 5.0),
    shuffle=True,
    random_state=None,
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N_FEATURES)])
df["y"] = data[1]
figure

By decreasing cluster_std we make them less separable.

data = make_blobs(
    n_samples=60,
    n_features=N_FEATURES,
    centers=3,
    cluster_std=2.5,
    center_box=(-5.0, 5.0),
    shuffle=True,
    random_state=None,
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(N_FEATURES)])
df["y"] = data[1]
figure

Anisotropic data

data = make_blobs(n_samples=50, n_features=2, centers=3, cluster_std=1.5)
transformation = [[0.5, -0.5], [-0.4, 0.8]]
data_0 = np.dot(data[0], transformation)
df = pd.DataFrame(data_0, columns=[f"f{i}" for i in range(1, 3)])
df["y"] = data[1]
figure

Concentric clusters

Sometimes we might be interested in creating a non-separable cluster. The simplest way is to create concentric clusters with the make_circles method.

from sklearn.datasets import make_circles

data = make_circles(
    n_samples=N_SAMPLES, shuffle=True, noise=None, random_state=random_state, factor=0.6
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(2)])
df["y"] = data[1]
figure

Adding noise

The noise parameter allows to create a concentric noisy dataset.

data = make_circles(
    n_samples=N_SAMPLES, shuffle=True, noise=0.15, random_state=random_state, factor=0.6
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(2)])
df["y"] = data[1]
figure

Moon clusters

A shape that can be useful to other methods (such as Counterfactuals, for instance) is the one generated by the make_moons method.

from sklearn.datasets import make_moons

data = make_moons(
    n_samples=N_SAMPLES, shuffle=True, noise=None, random_state=random_state
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(2)])
df["y"] = data[1]
figure

Adding noise

As usual, the noise parameter allows to control the noise.

data = make_moons(
    n_samples=N_SAMPLES, shuffle=True, noise=0.1, random_state=random_state
)
df = pd.DataFrame(data[0], columns=[f"f{i+1}" for i in range(2)])
df["y"] = data[1]
figure

Generative AI for synthetic data

For more advanced synthetic data generation using deep learning and generative models, please see:

These methods learn the underlying data distribution from real datasets and generate new synthetic samples that preserve statistical properties and relationships between features.

Text generation

For generating synthetic text data, Large Language Models (LLMs) are commonly used:

These approaches can generate realistic text that maintains the style, tone, and domain characteristics of the training data, useful for privacy-preserving data sharing and increasing dataset diversity.

Time-series data

Random walk

See [Random walk].

Simple periodic

No trend

Generate a simple HMM with a sine state and gaussian observations:

import numpy as np

def generate_sine(period, n):
    cycles = n / period
    length = np.pi * 2 * cycles
    return np.sin(np.arange(0, length, length / n))

We will now get a set of n=1000n=1000 observations with a p=10p=10 period

N=1000
data = generate_sine(10, N) * np.random.uniform(10, size=N)
figure

Trend

N=1000
data = (generate_sine(10, N) * np.random.uniform(10, size=N)) + np.arange(N)/200.0
figure

Univariate data

Using the streamad library:

from streamad.util.dataset import CustomDS
from streamad.util import StreamGenerator, plot

ds = CustomDS("_data/streamad/uniDS.csv")

stream = StreamGenerator(ds.data)
figure