如何将 Tensorflow 与 Estimator 结合使用,以使用增强树显示数据样本?

tensorflowserver side programmingprogramming更新于 2025/5/20 4:52:17

可以使用 Tensorflow 的增强树通过‘head’方法、‘describe’方法和‘shape’方法显示 titanic 数据集的样本。head 方法提供数据集的前几行,而 describe 方法提供有关数据集的信息,例如列名、类型、平均值、方差、标准差等。 shape 方法给出了数据的维度。

阅读更多: 什么是 TensorFlow,以及 Keras 如何与 TensorFlow 配合使用来创建神经网络?

我们将使用 Keras Sequential API,它有助于构建用于处理普通层堆栈的顺序模型,其中每个层都有一个输入张量和一个输出张量。

包含至少一个层的神经网络称为卷积层。我们可以使用卷积神经网络构建学习模型。

我们使用 Google Colaboratory 运行以下代码。Google Colab 或 Colaboratory 有助于在浏览器上运行 Python 代码,无需配置,可以免费访问 GPU(图形处理单元)。Colaboratory 建立在 Jupyter Notebook 之上。

我们将了解如何使用决策树和 tf.estimator API 训练梯度提升模型。

Estimator 是 TensorFlow 对完整模型的高级表示。它旨在轻松扩展和异步训练。估算器使用特征列来描述模型如何解释原始输入特征。估算器需要一个数字输入向量,特征列将有助于描述模型应如何转换数据集中的每个特征。

提升树模型被认为是回归和分类最流行和最有效的机器学习方法。它是一种集成技术,结合了许多(10 个、100 个或 1000 个)树模型的预测。

示例

print("Some sample of the data")
print(dftrain.head())
print("Metadata about the dataset")
print(dftrain.describe())
print("Dimensions of the data")
print(dftrain.shape[0], dfeval.shape[0])

代码来源 −https://www.tensorflow.org/tutorials/estimator/boosted_trees

输出

Some sample of the data
   sex    age   n_siblings_spouses parch ... class deck embark_town     alone
0  male   22.0   1                 0    ... Third unknown Southampton    n
1  female 38.0   1                 0    ... First C Cherbourg            n
2  female 26.0   0                 0    ... Third unknown Southampton    y
3  female 35.0   1                 0    ... First C Southampton          n
4  male   28.0   0                 0    ... Third unknown Queenstown     y
[5 rows x 9 columns]
Metadata about the dataset
        age       n_siblings_spouses parch fare
count   627.000000 627.000000 627.000000 627.000000
mean    29.631308   0.545455   0.379585    34.385399
std     12.511818   1.151090   0.792999    54.597730
min     0.750000    0.000000   0.000000    0.000000
25%     23.000000   0.000000   0.000000    7.895800
50%     28.000000   0.000000   0.000000    15.045800
75%     35.000000   1.000000   0.000000   31.387500
max     80.000000   8.000000   5.000000   512.329200
Dimensions of the data
627 264

解释

  • 数据集包含一个训练集和一个评估集。
  • dftrain 和 y_train 是训练集。
  • 它被模型用来学习特征和模式。
  • 使用评估集、dfeval 和 y_eval 测试模型。
  • 获取数据的某些汇总统计数据,并显示在控制台上。

相关文章