如何使用 Tensorflow 和 Estimator 定义数据集训练和评估的输入函数?

tensorflowserver side programmingprogramming更新于 2025/5/20 2:37:17

Tensorflow 和 Estimator 可用于定义数据集训练和评估的输入函数,该函数使用特征和标签生成字典。这是使用‘from_tensor_slices’方法实现的。此函数还将对数据集中的数据进行混洗,并定义训练步骤的数量。最后,此函数返回有关数据集的组合数据作为输出。通过将训练数据集传递给此函数来调用它。

阅读更多: 什么是 TensorFlow,以及 Keras 如何与 TensorFlow 配合使用来创建神经网络?

我们将使用 Keras Sequential API,它有助于构建用于处理普通层堆栈的顺序模型,其中每个层都有一个输入张量和一个输出张量。

包含至少一个层的神经网络称为卷积层。我们可以使用卷积神经网络构建学习模型。

我们使用 Google Colaboratory 运行以下代码。Google Colab 或 Colaboratory 有助于在浏览器上运行 Python 代码,无需任何配置,可以免费访问 GPU(图形处理单元)。Colaboratory 建立在 Jupyter Notebook 之上。

让我们了解如何使用估算器。

估算器是 TensorFlow 对完整模型的高级表示。它专为轻松扩展和异步训练而设计。

我们将使用 tf.estimator API 训练逻辑回归模型。该模型用作其他算法的基线。我们使用泰坦尼克号数据集,目标是预测乘客的生存情况,给定性别、年龄、等级等特征。

估算器使用特征列来描述模型如何解释原始输入特征。估算器需要一个数字输入向量,特征列将有助于描述模型应如何转换数据集中的每个特征。

示例

print("The entire batch of dataset is used since it is small")
NUM_EXAMPLES = len(y_train)
print("Function that iterates through the dataset, in memory training")
def make_input_fn(X, y, n_epochs=None, shuffle=True):
def input_fn():
dataset = tf.data.Dataset.from_tensor_slices((dict(X), y))
if shuffle:
   dataset = dataset.shuffle(NUM_EXAMPLES)
   dataset = dataset.repeat(n_epochs)
   dataset = dataset.batch(NUM_EXAMPLES)
   return dataset
return input_fn
print("Training and evaluation input function have been defined")
train_input_fn = make_input_fn(dftrain, y_train)
eval_input_fn = make_input_fn(dfeval, y_eval, shuffle=False, n_epochs=1)

代码来源 −https://www.tensorflow.org/tutorials/estimator/boosted_trees

输出

The entire batch of dataset is used since it is small
Function that iterates through the dataset, in memory training
Training and evaluation input function have been defined

解释

  • 需要创建输入函数。
  • 它指定如何将数据读入模型进行训练和推理。
  • tf.data API 中的 from_tensor_slices 方法用于直接从 Pandas 库读取数据。
  • 这适用于较小的内存数据集。
  • 如果数据集很大,则可以使用 tf.data API,因为它支持多种文件格式。

相关文章